Lepton

Lepton AI's "Build AI The Simple Way" tool is a developer-centric platform that allows users to efficiently run AI applications at scale in...

Last verified:

Visit Lepton

What is Lepton?

NVIDIA DGX Cloud Lepton (formerly Lepton AI) is an AI platform that connects developers to global GPU compute across a network of cloud providers. It provides a unified experience for development, training, and inference, allowing teams to build, train, and deploy AI applications without managing underlying infrastructure. The platform bridges the gap between AI demand and global GPU supply by bringing together NVIDIA Cloud Partners, GPU marketplaces, cloud providers, and local environments in a single developer-friendly platform.

Key features include frictionless deployment across any GPU cloud with minimal operational burden, the ability to run compute in specific regions for data sovereignty compliance and low-latency requirements, and a unified experience across development, training, and inferencing. The platform offers instant access to NVIDIA's accelerated APIs including serverless endpoints, prebuilt NVIDIA NIM microservices, and GPU-backed compute. It supports distributed training jobs, batch processing, managed endpoints, and dev sessions with managed GPUs.

DGX Cloud Lepton is designed for AI natives, model builders, and teams that iterate quickly. It is suitable for AI developers building and deploying machine learning models, research teams working on AI development projects, and tech companies implementing AI solutions at scale. The platform is enterprise-compliant with SOC2 and HIPAA, maintaining 100% uptime while processing over 20 billion tokens and generating over 1 million images daily.

Lepton pricing

Pricing model: Free

Lepton AI offers three pricing plans: Basic, Standard, and Enterprise. The Basic Plan has no subscription fees and is pay-per-use, allowing up to 4 CPUs, 16 GB memory, and 1 GPU cumulatively. New users receive $10 free credits with a 10 RPM rate limit and 10M tokens/day free tier. The Standard Plan is $30 per month plus usage-based compute charges, offering more resources with dedicated support and advanced features. The Enterprise Plan offers tailored pricing for large-scale deployments with advanced customization, unlimited user seats, dedicated account management, and 24/7 priority support. Compute is billed by minute, storage is $0.153/GB/month (free under 1GB), and network traffic is free for first 10GB/month then $0.15/GB. Model API pricing includes Llama 3 70B at $0.80 per 1M tokens (input/output), Llama 3 8B at $0.07 per 1M tokens, and dedicated endpoints priced per GPU-hour for H100/A100 classes.

Lepton pros

  • Python-native development without containers or Kubernetes
  • Single-command deployment from local to cloud
  • Unified experience across development, training, and inference
  • Access to global GPU network across multiple cloud providers
  • Frictionless multi-cloud AI deployment
  • Run compute in specific regions for data sovereignty
  • Instant access to NVIDIA accelerated APIs and NIM microservices
  • Distributed training job support
  • Batch processing capabilities
  • Managed endpoints with unique URLs per workload
  • Dev sessions with managed GPUs
  • Enterprise-grade performance and reliability
  • SOC2 and HIPAA compliance
  • 100% uptime guarantee
  • Automatic scaling and resource optimization
  • OpenAPI-compatible chat surface
  • No rearchitecting when compute changes
  • Prototype to production faster
  • Predictable enterprise-grade performance
  • GPU marketplace integration

Lepton cons

  • Basic plan limited to 4 CPUs, 16 GB memory, 1 GPU
  • 10 requests per minute rate limit on basic plan
  • Storage costs $0.153 per GB per month after 1GB
  • Network egress charged at $0.15/GB after first 10GB free
  • Standard plan requires $30 monthly subscription plus usage
  • New users only get $10 free credits
  • Learning curve for new users to understand platform architecture
  • Limited free tier with restricted feature access
  • Enterprise plan can be expensive for large deployments
  • No JavaScript client yet (Python and cURL only)

Frequently asked questions about Lepton

What is NVIDIA DGX Cloud Lepton?

NVIDIA DGX Cloud Lepton is an AI platform that connects developers to global GPU compute across a network of cloud providers. It brings together a global network of NVIDIA Cloud Partners, GPU marketplaces, cloud providers, and local environments to streamline discovery, development, and deployment of AI workloads in a single, developer-friendly platform with unified experience for development, training, and inference.

Who is DGX Cloud Lepton designed for?

DGX Cloud Lepton is designed for AI natives, model builders, and teams that iterate quickly. It is suitable for AI developers building and deploying machine learning models, research teams working on AI research and development projects, and tech companies looking to implement AI solutions at scale.

How do I get started with Lepton AI?

To get started, install the Python SDK with pip install -U leptonai, create a workspace via the Lepton AI dashboard at https://dashboard.lepton.ai/, then login using lep login -c <your_workspace_credential>. You can find workspace credentials in the Settings-Tokens page. After login, verify access with lep workspace list and begin deploying models or calling APIs.

What models are available on Lepton AI?

Lepton AI offers 14 tracked models including Llama 3.1 70B, Llama 3 70B Instruct, Llama 3 8B Instruct, Llama 3.3 70B with 128k context, Mixtral 8x7B with 32k context, Dolphin 2.6 Mixtral 8x7B, MythoMax L2 13B, OpenChat 3.5, and WhisperX/Whisper Large v3 for audio processing.

How do I deploy an LLM model on Lepton?

You can deploy an LLM using the Lepton CLI: create a photon with lep photon create -n myphoton -m hf:gpt2, push it with lep photon push -n myphoton, then run it with lep photon run -n myphoton. This creates a deployment with endpoints you can interact with via Python client, cURL, or Web UI.

What are the pricing plans available?

Lepton offers three plans: Basic (no subscription, pay-per-use, up to 4 CPUs/16GB/1 GPU), Standard ($30/month plus usage with more resources and dedicated support), and Enterprise (tailored pricing with unlimited seats, customization, dedicated account management, and 24/7 priority support).

How do I interact with a deployed model?

You can interact with deployments via Python client (import Client from leptonai.client, create with workspace ID, deployment name, and token), raw cURL calls to RESTful endpoints, or Web UI at https://dashboard.lepton.ai/ which provides prefilled code examples. The Web UI also shows OpenAPI specs under the API tab.

What is the free tier for new users?

New users get $10 free credits, a rate limit of 10 requests per minute (RPM), and 10M tokens per day for free users. The basic plan has no subscription fee and you only pay for consumed resources, with the first 10GB network traffic per month free.

How is compute and storage billed?

Compute usage is billed by minute for specific resources. Storage is charged $0.153 per GB per month, but files under 1GB are free. Network ingress/egress is free for the first 10GB per month per workspace, then $0.15/GB for additional traffic.

What security and compliance features does Lepton offer?

Lepton AI is compliant with SOC2 and HIPAA for enterprise use, ensuring reliability through high availability and efficient compute features. The platform maintains 100% uptime and processes over 20 billion tokens daily while generating over 1 million images, with enterprise-grade performance, reliability, and security through cloud partners in the marketplace.

Categories

Use cases

Browse all AI tools on NeedAnAI