Novita

novita.ai is a platform that offers hundreds of fast and affordable AI image generation application programming interfaces (APIs). It boast...

Last verified:

Visit Novita

What is Novita?

Novita is an AI-native cloud platform that provides production-ready model inference, GPU infrastructure, and agent runtimes under a single API, designed for builders who need to run text, image, audio, and video models at scale. The platform exposes 200+ models via a unified API and supports serverless model calls billed by token, dedicated endpoints for guaranteed performance, and private GPU instances for predictable latency. Novita also offers serverless GPU jobs, full-control GPU instances, and bare-metal clusters so teams can choose between pay-as-you-go execution and isolated hardware for high-throughput or sensitive workloads. The product targets developers, ML engineers, startups, and enterprises that want to avoid managing GPU infrastructure while running large open-source models in production.

Novita pricing

Pricing model: Paid

Novita advertises a pay-as-you-go pricing model with token-based billing for model input and output (examples shown as per-MT input/output rates per model), a free-to-start experience (signup grants starting access/quotas), and higher-tier offerings such as private dedicated endpoints, full-control GPU instances, and bare-metal clusters for enterprise customers. Model pages list specific in/out MTokens pricing per model (for example, sample LLMs with distinct input/output MT token rates), while GPU Instances and serverless jobs are billed based on execution and instance selection; enterprise plans include dedicated support and SLA-backed private endpoints. Explicit per-tier free-plan quotas or an itemized monthly subscription table are not shown prominently on the public site.

Novita pros

  • Single unified API for 200+ models
  • Serverless model inference billed by tokens, not hours
  • Private dedicated endpoints for consistent latency
  • Full-control GPU instances provisioned in seconds
  • Serverless GPU jobs that scale to zero automatically
  • Bare-metal clusters for maximum throughput and low overhead
  • Large context windows offered on multiple models
  • Support for text, image, audio, and video workloads
  • Ability to add or bring your own models
  • Competitive price-performance claims versus major clouds
  • High-performance GPU hardware (H100/H200) available
  • Isolated agent sandbox runtimes for secure automation
  • Production-focused features (SLA-oriented endpoints)
  • Integrated tooling for agents and runtime orchestration
  • Dedicated technical support for enterprise customers
  • Multiple deployment options (serverless, instances, bare metal)

Novita cons

  • Pricing presented per MToken metrics which can be unclear for newcomers
  • No public detailed free-tier quotas shown on main pages
  • Some advanced features (bare metal, private endpoints) aimed at enterprises
  • Billing by token may complicate cost predictions for multimodal workloads
  • Documentation references (llms.txt) require extra navigation to find full model lists
  • Not all model performance or latency SLAs are listed per-region on public pages
  • Feature details for image/video pipelines are summarized but not fully itemized
  • Onboarding for self-hosted model uploads and custom training has limited public detail

Frequently asked questions about Novita

How do I start using Novita and call models via the API?

Sign up on Novita, obtain an API key, and use the unified Novita API or official SDKs to call any of the 200+ available models; quickstart guides in the docs walk through authentication, example requests, and code snippets for common languages.

What billing model does Novita use for model inference?

Novita bills inference by tokens (MTokens) for model input and output with per-model input/output rates displayed on model listings, and for GPU jobs you pay for execution time or chosen instance types, while serverless GPU jobs scale to zero when idle to avoid idle-instance charges.

Can I get guaranteed performance for production workloads?

Yes — Novita offers private dedicated endpoints and isolated GPU instances so you get consistent latency and throughput without noisy neighbors, which is intended for production deployments that require predictable performance.

Does Novita support running custom or third-party models?

Novita allows adding or bringing your own models and exposes many community/open-source models (including integrations with model sources), with options to deploy them on serverless inference, dedicated instances, or bare-metal clusters depending on performance needs.

What GPU hardware and cluster options are available?

Novita provides multiple options including full-control GPU instances (H100/H200), serverless GPUs for on-demand jobs, and bare-metal clusters with NVLink and high-bandwidth RDMA interconnects for maximum throughput and minimal abstraction overhead.

How does Novita handle security and isolation for agent runtimes?

Agent Sandbox runtimes are purpose-built, isolated environments (not user-configured containers) that securely run AI-generated code, tools, and model calls with isolation to prevent cross-agent interference and to support reproducible agent execution.

Is there a free tier or trial to test the platform?

The site states Novita is 'free to start' and documentation/marketing references signup and starter quotas, but explicit free-tier quotas and detailed trial limits are not prominently listed on the main pages and typically require signing up to view account-specific credits.

How do I choose between serverless GPUs and dedicated instances?

Choose serverless GPUs for job-based, bursty workloads where you want automatic scaling and pay only for execution, while dedicated instances or bare-metal clusters are better for sustained, latency-sensitive, or high-throughput production workloads that need predictable performance.

What SLAs or support levels does Novita provide for enterprises?

Novita advertises dedicated support and enterprise-focused offerings (private endpoints, bare metal) and highlights fast technical support, but specific SLA documents and enterprise contract details are provided during sales/enterprise onboarding rather than on the public website.

Can I run multimodal tasks (text, image, audio, video) in one workflow?

Yes — Novita's platform supports multimodal workloads with APIs and serverless runtimes that run text, image, audio, and video models through the same platform and unified API, enabling combined pipelines and agent-driven orchestration.

Categories

Use cases

Browse all AI tools on NeedAnAI