Twelve Labs

Twelve Labs is an AI-driven video search platform that helps developers create applications that can see, listen, and understand the world ...

Last verified:

Visit Twelve Labs

What is Twelve Labs?

Twelve Labs is an AI-powered video understanding platform that transforms raw video into searchable, AI-ready data at massive scale. The platform uses state-of-the-art multimodal foundation models (Marengo for embeddings and Pegasus for generative analysis) to identify objects, actions, speech, and text in video content. It equips developers with deep semantic search capabilities, allowing users to find exact moments within videos using natural language queries instead of tags or metadata.

Key features include semantic search across speech, text, audio, and visuals; dynamic video-to-text generation for summaries, chapters, and highlights; image-to-video search; multimodal embeddings for cross-modal searches; automatic content segmentation; compliance scanning for policy risks and brand safety; highlight creation; and insight generation at scale. The platform indexes an hour of video in a minute, processing up to 10k+ hours per day at ~60x real-time speed.

Twelve Labs is designed for organizations working with video at scale, including teams in media, sports, advertising, government, security, and creative industries. It is particularly valuable for content creators, media management teams, post-production professionals, AdTech companies needing contextual targeting, public sector organizations for evidence management, and developers building semantic search, recommendations, or video analysis applications.

Twelve Labs pricing

Pricing model: Freemium

Play for free, pay as you go. Three tiers available: For testing and building includes indexing limit of less than 10 hours, shared environment, no org account, no SSO/SAML, and no finetune. For launching and growing includes unlimited indexing, shared environment, no org account, no SSO/SAML, and no finetune. For scaling and services includes unlimited indexing, dedicated environment, included org account, included SSO/SAML, and included finetune. API keys are obtained from the dashboard page at playground.twelvelabs.io.

Twelve Labs pros

  • Deep semantic search with natural language queries instead of tags
  • Indexes an hour of video in just one minute
  • Processes up to 10k+ hours per day at ~60x real-time speed
  • Multimodal approach searching across speech, text, audio, and visuals
  • Image-to-video search for finding videos similar to provided images
  • One-time video indexing for multiple tasks and repurposing
  • Pegasus 1.5 model achieves #1 on Video-MME benchmark
  • Marengo model achieves 78.5% composite accuracy across 47 languages
  • Generates summaries, chapters, highlights, and custom reports from videos
  • Automatic content segmentation identifying natural breaks and scene changes
  • Compliance scanning for policy risks and brand safety issues
  • SOC 2 Type II certified with encrypted data handling
  • Flexible deployment: on-premise, hybrid, or cloud-based environments
  • Fine-tuning capabilities available with few examples for better results
  • Simple API integration with just a few API calls
  • Rapid result retrieval within seconds
  • Cloud-native distributed infrastructure handling thousands of concurrent requests
  • Pegasus 1.5 outperforms Gemini 3.1 Pro on multimodal prompting by 13.1%
  • 10x faster content review and compliance scanning
  • Trusted by NFL Media, Sejong City, and MindsDB

Twelve Labs cons

  • Fine-tuning requires contacting sales team at [email protected]
  • Dedicated environment and SSO/SAML only available on highest tier
  • Free tier limited to less than 10 hours indexing
  • Shared environment on testing and launching tiers
  • Pegasus model supports videos up to only two hours maximum
  • No public pricing numbers disclosed on website
  • Pay-as-you-go model may be unpredictable for budget planning
  • Requires API key authentication for all requests
  • Limited to visual and audio modalities currently

Frequently asked questions about Twelve Labs

What is Twelve Labs and what does it do?

Twelve Labs is a video understanding platform that uses multimodal foundation models to identify objects, actions, speech, and text in video content. It transforms raw video into searchable, AI-ready data, enabling deep semantic search, video-to-text generation, and multimodal embeddings at massive scale.

What are the Marengo and Pegasus models?

Marengo is the multimodal embedding model that turns video into spatiotemporal embeddings for search and embedding tasks, achieving 78.5% composite accuracy across 47 languages. Pegasus is the video language model that reasons continuously over full temporal arcs of videos up to two hours, tracking entities, causation, and narrative, and is #1 on Video-MME benchmark.

How fast does Twelve Labs process video?

The platform indexes an hour of video in one minute and processes up to 10k+ hours per day at approximately 60x real-time speed through a single multimodal pipeline.

Can I search videos using images?

Yes, Twelve Labs supports image-to-video search, allowing you to perform searches using images as queries and find videos semantically similar to the provided images. This addresses challenges when reverse image search tools yield inconsistent results or when describing desired results using text is difficult.

What industries use Twelve Labs?

Twelve Labs is built for teams in media, sports, advertising, government, security, and creative industries. Specific use cases include NFL Media for content packaging, Sejong City for public sector technology deployment, and AdTech companies for contextual targeting.

Is Twelve Labs secure and compliant?

Yes, Twelve Labs is SOC 2 Type II certified with encrypted data handling. The entire intelligence stack deploys where customers want, including on-premise, hybrid, or cloud-based environments.

How do I get started with Twelve Labs?

You can get your API keys from the dashboard page at playground.twelvelabs.io/dashboard/api-key. The platform offers SDKs for Python and Node.js, and you can test capabilities in Playground before writing any code. Documentation, sample apps, and quick-start tutorials are available at the Developer Hub.

What video length does Pegasus support?

Pegasus reasons continuously over the full temporal arc of any video asset up to two hours, tracking entities, causation, and narrative across time. It is not a transcript reader but a video reasoner.

Can I fine-tune Twelve Labs models?

Yes, fine-tuning capabilities are available to help achieve better results with only a few examples. However, for details on fine-tuning the models, you must contact the sales team at [email protected].

What makes Twelve Labs different from other video AI solutions?

Twelve Labs uses a video-first, multimodal approach surpassing traditional unimodal models that depend exclusively on text or images. It offers simplified API integration with rich tasks in few calls, natural language use instead of tags/keywords, one-time video indexing for multiple tasks, image-to-video search, flexible deployment options, and fine-tuning capabilities.

Categories

Use cases

Browse all AI tools on NeedAnAI