Ray

Ray is an AI compute engine. Ray consists of a core distributed runtime and a set of AI Libraries for accelerating ML workloads.

Last verified:

Visit Ray

What is Ray?

Ray is an open-source AI compute engine and unified distributed computing framework designed to scale AI and Python workloads from a laptop to thousands of GPUs. It precisely orchestrates infrastructure for any distributed workload on any accelerator at any scale, solving the AI Complexity Wall that teams face with complex data modalities and new models.

Key features include Ray Core for scaling Python code with primitives like tasks, actors, and objects; Ray Libraries for end-to-end AI solutions including model training, model serving with Ray Serve, batch inference, reinforcement learning with RLlib, Gen AI workflows, LLM inference, and LLM fine-tuning. It supports multi-modal data processing (images, videos, audio), parallel Python code for simulation and backtesting, and is compatible with any ML framework.

Ray is built for AI/ML engineers, data scientists, and developers who need to scale distributed AI workloads. It serves industry leaders building foundation models with 300B+ parameters, teams needing 1M+ CPU cores for online model serving, and organizations wanting to reduce data processing costs by 82%.

The ecosystem includes Ray Core for distributed Python applications, high-level Ray Libraries for ML workloads, and tools for deploying Ray clusters, debugging, and optimizing applications. It runs on any machine, cluster, cloud provider, and Kubernetes with 34.8k GitHub stars and 1000+ contributors.

Ray pricing

Pricing model: Freemium

Ray itself is open-source and free to use under Apache-2.0 license. Anyscale offers a fully managed AI platform for Ray with $100 starter credits to get started. Anyscale uses usage-based pay-as-you-go billing with no monthly fixed fees. Pricing includes: CPU Only AC $0.0135/hr, NVIDIA T4 AC $0.5682/hr, NVIDIA L4 AC $0.9542/hr, NVIDIA A10G AC $1.3635/hr, NVIDIA A100 AC $4.9591/hr, NVIDIA H100 AC $9.2880/hr, NVIDIA H200 AC $10.6812/hr. Two deployment options: Hosted (fully managed, limited regions, business hours support) and Bring Your Own Cloud (any cloud/region, on-prem, 24x7 enterprise SLAs). Volume discounts unlock as usage grows. Committed contracts available for additional discounts.

Ray pros

  • Python-native framework built by developers for developers
  • Scale seamlessly from laptop to thousands of GPUs with no code changes
  • Supports any AI or ML workload including Gen AI and traditional ML
  • Handles any data types and model architectures
  • Uses heterogeneous GPUs and CPUs with fine-grained independent scaling
  • Fully utilizes every accelerator for maximum efficiency
  • Ray Serve offers independent scaling and fractional resources
  • Distributed training with 1 line of code for any framework
  • Ray RLlib supports production-level distributed RL workloads
  • 82% lower data processing cost ($120M savings per year documented)
  • 30x cost reduction switching from Spark to Ray for batch inference with GPUs
  • 4x improvement in GPU utilization and 7X lower costs
  • 10-100x more model training data capability
  • 10x more models trained per month documented
  • 12x faster iteration to deliver 100+ production models
  • Open source with 40k+ GitHub downloads and active community
  • Runs on any cloud provider (AWS, Azure, GCP) and on-premises
  • Supports LLM inference (online and batch) and fine-tuning at scale
  • Multi-modal data processing for structured and unstructured data

Ray cons

  • Steep learning curve for distributed computing concepts
  • Complex setup and cluster management without managed platform
  • Requires significant infrastructure knowledge for production deployment
  • Debugging distributed applications can be challenging
  • Memory management requires careful optimization
  • Community support primarily through Slack (not enterprise SLA)
  • Open source version lacks enterprise governance features
  • GPU costs can be expensive without careful optimization
  • Integration with existing pipelines requires code modifications

Frequently asked questions about Ray

What is Ray?

Ray is an open-source AI compute engine and unified distributed computing framework that makes it easy to scale AI and Python workloads from a laptop to a cluster. It precisely orchestrates infrastructure for any distributed workload on any accelerator at any scale, supporting any AI or ML workload, any data types and model architectures, with heterogeneous GPUs and CPUs.

Is Ray free to use?

Yes, Ray is open-source under the Apache-2.0 license and free to use. You can download it from GitHub with 40k+ repo downloads. Anyscale offers a fully managed platform with $100 starter credits, but the core Ray framework is free and can run on any machine, cluster, cloud provider, or Kubernetes.

What workloads does Ray support?

Ray supports any AI or ML workload including parallel Python code, multi-modal data processing (images, videos, audio), model training (Gen AI foundation models, time series, XGBoost), model serving with Ray Serve, batch inference, reinforcement learning with RLlib, Gen AI workflows, LLM inference (online and batch), and LLM fine-tuning at scale.

How does Ray scale compared to traditional frameworks?

Ray scales seamlessly from your laptop to thousands of GPUs with no code changes. Documented results show 10-100x more model training data, 30x cost reduction switching from Spark for batch inference with GPUs, 4x improvement in GPU utilization, 10x more models trained per month, and 12x faster iteration for production models.

What is Ray Core?

Ray Core provides a small number of core primitives (tasks, actors, objects) for building and scaling distributed Python applications. It is Python-native and allows you to scale and distribute any Python code for use cases like simulation, backtesting, and more with a simple API.

What is Ray Serve?

Ray Serve is a model serving library that lets you deploy models and business logic, not instances. It offers independent scaling and fractional resources to maximize deployed models. It supports any ML model from LLMs to stable diffusion models to object detection models with fine-grained scaling.

Can Ray run on my cloud provider?

Yes, Ray runs on any machine, cluster, cloud provider, and Kubernetes. Anyscale supports deployment on AWS, Azure, GCP, Nebius, or CoreWeave. You can use the Hosted option (Anyscale-managed, limited regions) or Bring Your Own Cloud option (any cloud/region, on-prem, your VPC).

What is the Ray community like?

Ray has a growing open-source community with 34.8k GitHub stars, 40k+ repo downloads, and 1000+ contributors. You can join the Ray Community Slack to connect with fellow AI practitioners, get help, and stay in the loop with updates.

How does Anyscale differ from open-source Ray?

Anyscale is a fully managed AI platform for Ray built by the team that created Ray. It includes all benefits of Ray plus enterprise governance, advanced developer tooling, expert support, 24x7 coverage with enterprise SLAs, unlimited case submissions, and deployment in your own VPC for data residency.

What results can I expect from using Ray?

Industry leaders report 10-100x more model training data, 1M+ CPU cores deployed for online serving, 300B+ parameters for foundation model training, 82% lower data processing costs ($120M/year savings), 30x cost reduction from Spark, 4x GPU utilization improvement with 7X lower costs, 10x more models monthly, and 12x faster iteration for 100+ production models.

Categories

Browse all AI tools on NeedAnAI