Together AI

Build, deploy, and optimize AI models with ultra-fast, scalable solutions.. [Contact for Pricing]

Last verified:

Visit Together AI

What is Together AI?

Together AI is a full-stack AI Native Cloud platform that empowers developers, researchers, and organizations to train, fine-tune, and deploy generative AI models. The platform accelerates inference by 2x through cutting-edge research, reduces costs by 60% with workload-specific optimization, and speeds up pre-training by 90% using the Together Kernel Collection. It serves as a research-driven, open-source AI company building technology to benefit society.

Key features include Serverless Inference for running 50+ leading open-source models on demand without infrastructure management, Batch Inference for cost-effectively processing massive workloads up to 30 billion tokens per model, and Dedicated Model Inference for teams needing speed and control. The platform offers Accelerated Compute scaling from instant clusters to thousands of GPUs, Fine-Tuning capabilities using LoRA and full fine-tuning, managed storage with zero egress fees, and GPU clusters for large-scale training. Together AI also provides Sandboxes for AI app development and supports multimodal workloads including image, video, audio, and vision models.

The platform is designed for developers building AI applications, researchers conducting AI systems research, and enterprises needing production-grade AI infrastructure. Customers like Cursor use Together AI for real-time low-latency inference at scale, while Decagon engineered sub-second voice AI with 6x cost reduction and 11x faster inference. The platform supports chat models, vision models, video generation, text-to-speech, embedding models, rerank models, and code models from organizations like Meta (Llama), Qwen, DeepSeek, Mistral AI, Google, and OpenAI.

Together AI is grounded in cutting-edge systems research, contributing leading open-source models and datasets including FlashAttention, ThunderKittens, RedPajama-v2 (30 trillion token dataset), and Mamba-3. The platform enables function calling, JSON mode, structured outputs, multi-turn conversations, and audio transcription through a unified API with official SDKs for Python and TypeScript.

Together AI pricing

Pricing model: Freemium

Serverless Inference: Pay per 1M tokens with free tier available for Llama 3.3 70B Instruct Turbo. Prices range from $0.05/1M tokens (gpt-oss-20B) to $2.10/1M tokens (DeepSeek V4 Pro). Image models priced per image ($0.0006-$0.08 depending on model). Video models priced per video ($0.14-$3.20). Audio transcription at $0.0015 per minute (Whisper Large v3). Dedicated Inference: 1x H100 80GB at $6.49/hour, 1x HGX B200 180GB at $11.95/hour. GPU Clusters: On-demand NVIDIA HGX H100 at $5.49/hour, reserved at $4.99/hour. Fine-tuning: $0.48/1M tokens for models up to 16B, $1.50 for 17B-69B, $2.90 for 70-100B, up to $40/1M tokens for GLM-5. Sandbox: $0.0446 per vCPU, $0.0149 per GiB RAM. Storage: $0.16/GiB/month with zero egress fees.

Together AI pros

  • 2x faster inference powered by cutting-edge research
  • 60% lower cost with workload-specific optimization
  • 90% faster pre-training with Together Kernel Collection
  • 50+ leading open-source models available via API
  • Serverless inference with no infrastructure to manage
  • No long-term commitments required
  • Batch inference scales to 30 billion tokens per model
  • Dedicated infrastructure for speed and control
  • Official SDKs for Python and TypeScript
  • REST API works from any language
  • Support for function calling and JSON mode
  • Fine-tuning with LoRA and full fine-tuning options
  • Zero egress fees on managed storage
  • GPU clusters scale from instant to thousands of GPUs
  • Free tier available for Llama 3.3 70B Instruct Turbo
  • Support for multimodal models (image, video, audio, vision)
  • Long context support up to 1M tokens (Llama 4 Maverick)
  • Private deployment options available
  • Cutting-edge research contributions (FlashAttention, ThunderKittens)
  • Open-source model contributions (Mamba-3, RedPajama-v2)

Together AI cons

  • Free tier has reduced rate limits (.6 requests/minute)
  • Image models require credits (cannot use zero balance)
  • Free Llama 3.3 70B has only 8192 context vs 131072 paid
  • Some models are deprecated (check deprecation notices)
  • Pricing varies by model resolution/duration settings
  • Dedicated H200 and B200 require contacting sales
  • Batch API prices significantly higher than standard
  • Fine-tuning requires labeled data (hundreds to thousands)
  • Not ideal for factual grounding (RAG recommended instead)
  • Quantized models (Turbo/Lite) have reduced precision
  • Rate limits vary by model (check model-specific limits)
  • FLUX.1 schnell free has 10 img/min rate limit
  • Video models limited to 5-10 second durations
  • Audio transcription streaming is expensive ($0.27/min)
  • Some premium models very expensive (GLM-5 at $40/1M tokens)
  • Dedicated endpoints require steady traffic for cost-effectiveness

Frequently asked questions about Together AI

What is Together AI?

Together AI is a full-stack AI Native Cloud platform that empowers developers and researchers to train, fine-tune, and deploy generative AI models. It is a research-driven, open-source AI company powered by cutting-edge systems research, offering faster inference (2x), lower costs (60%), and faster pre-training (90%) with the Together Kernel Collection.

How do I get started with Together AI?

Create an API key by registering at api.together.ai, then go to your project's API keys page to create and copy your key. Install the official SDK for Python (pip install together) or TypeScript (npm install together-ai), export your API key as TOGETHER_API_KEY environment variable, and make your first chat completion request with just a few lines of code.

What models are available on Together AI?

Together AI offers 50+ leading open-source models including Meta's Llama 4 Maverick/Scout, Llama 3.3 70B, Qwen3 models, DeepSeek-V3/R1, Mistral AI models, Google Gemma, and OpenAI's gpt-oss. Models span chat, vision, video, audio, code, embedding, and rerank categories with context lengths up to 1M tokens.

Is there a free tier available?

Yes, Together AI offers a free tier for Llama 3.3 70B Instruct Turbo (meta-llama/Llama-3.3-70B-Instruct-Turbo-Free) with a reduced rate limit of 0.6 requests/minute (36/hour) for free tier users. FLUX.1 [schnell] free also has a rate limit of 10 images/minute. Free tier models have reduced performance compared to paid Turbo endpoints.

What fine-tuning options does Together AI support?

Together supports LoRA (Low-Rank Adaptation) which is faster and cheaper, and full fine-tuning that updates every weight. Specialized fine-tuning includes preference fine-tuning (DPO), function-calling fine-tuning, reasoning fine-tuning, vision-language fine-tuning, and bring-your-own-model. Fine-tuning is billed per token of training data scaled by model size.

How does Together AI pricing work for image generation?

FLUX model pricing (except pro) is based on generated image size in megapixels (MP = Width × Height ÷ 1,000,000) and number of steps if exceeding default. Default steps vary by model (e.g., FLUX.1 [schnell] uses 4 steps, FLUX.1 Dev uses 28). Prices range from $0.0006 per image (Dreamshaper) to $0.08 (Qwen Image 2.0 Pro).

What is the difference between Serverless and Dedicated Inference?

Serverless Inference is the fastest way to run open-source models on demand with no infrastructure to manage and no long-term commitments, ideal for most teams starting out. Dedicated Model Inference deploys models on dedicated single-tenant GPU instances (H100 at $6.49/hour, B200 at $11.95/hour) for teams needing speed, control, and best economics at scale.

Can I use Together AI for enterprise workloads?

Yes, Together AI powers enterprise customers like Cursor delivering real-time low-latency inference at scale and Decagon engineering sub-second voice AI with 6x cost reduction. Enterprise features include dedicated endpoints, private deployments, GPU clusters scaling to thousands of GPUs, managed storage with zero egress fees, and dedicated infrastructure for speed and control.

What research does Together AI contribute?

Together AI contributes leading open-source research including FlashAttention-4, ThunderKittens optimized for NVIDIA Blackwell, Mamba-3 (open-source SSM), RedPajama-v2 (30 trillion token dataset), Aurora inference acceleration, Plan-divide-conquer for long context tasks, DeepCoder (14B coder at O3-mini level), Together MoA, and numerous kernel and inference optimization papers.

How do I handle streaming responses with Together AI?

To stream tokens as they arrive, add stream=True in Python or stream: true in TypeScript/cURL and iterate over response chunks. The same client works for multi-turn conversations, function calling, structured outputs, image generation, and audio transcription. Streaming is supported across all serverless inference models.

Categories

Use cases

Browse all AI tools on NeedAnAI