Speechmatics

Speechmatics is the world’s leading expert in Speech Intelligence, combining the latest breakthroughs in AI and ML to unlock the business value in human speech....

Last verified:

Visit Speechmatics

What is Speechmatics?

Speechmatics is a powerful speech-to-text (STT) and text-to-speech (TTS) API designed for enterprise-grade voice AI applications. It converts spoken language into accurate written text in 56+ languages and dialects, supporting real-time streaming and batch processing of pre-recorded audio files. The platform is built for real-world challenges including accents, background noise, and code-switching between languages, delivering 90%+ accuracy in diverse audio conditions.

Key features include sub-500ms latency for real-time transcription, speaker diarization that tracks multiple speakers even in overlapping conversations, word-level timestamps, code-switching support for bilingual models, and alphanumeric accuracy for phone numbers, postcodes, and account numbers (96.9% sequence accuracy on character strings). The API supports flexible deployment options including cloud (SaaS), on-premises, on-device, virtual appliance, and Docker containers. Additional capabilities include custom dictionary injection for up to 1,000 domain-specific terms, entity formatting for numbers/dates/currencies, profanity/hesitation detection, confidence scores per word, automatic language detection, and AI translation for over 30 language pairs.

Speechmatics is ideal for contact centers capturing account numbers and booking references, healthcare providers generating clinical notes with its Medical Model, media companies creating captions and summaries, conversational AI builders creating voice AI agents, court reporters and legal professionals needing real-time accuracy, and meeting platform developers automating note-taking. The platform offers ISO/IEC 27001:2022, GDPR, HIPAA, and SOC 2 Type II compliance for privacy-critical use cases.

Speechmatics pricing

Pricing model: Freemium

Speechmatics offers a free tier with 2,400 minutes (40 hours) per month of speech-to-text and 2 concurrent real-time sessions, plus 1 million characters (~20 hours) per month for text-to-speech. Pro tier pricing starts from $0.24 per hour of transcribed audio, with volume discounts automatically applied for usage above 500 hours per month (20% discount) and additional discounts available starting from 24,000 hours annually. Pro tier includes 50 concurrent real-time sessions, 10 file jobs per second, no rate limits, and access to all 56+ languages. Enterprise plans offer custom pricing, custom models, custom voice/language development, SaaS or on-premises deployment, highest concurrency, multi-region cloud options, and dedicated Customer Success Manager. The Startup Program provides up to $50,000 in credits with full API access and dedicated onboarding. Billing for Pro tier occurs monthly on the 1st for previous month usage; Enterprise billing is on a custom basis.

Speechmatics pros

  • 56+ languages and dialects with global coverage
  • 90%+ accuracy in real-world audio conditions
  • Sub-500ms latency for real-time transcription
  • Final transcripts in under 1 second, fastest real-time engine
  • Speaker diarization tracks multiple speakers in overlapping conversations
  • Code-switching support for bilingual conversations without accuracy loss
  • 96.9% sequence accuracy on character strings like phone numbers
  • Flexible deployment: cloud, on-prem, on-device, virtual appliance, containers
  • Custom dictionary for up to 1,000 domain-specific terms
  • Medical Model cutting errors on key terms by up to 50%
  • ISO/IEC 27001:2022, GDPR, HIPAA, and SOC 2 Type II compliant
  • AI translation for over 30 language pairs in single API call
  • Automatic language detection simplifies integration
  • Word-level timestamps for post-processing
  • Supports all major audio and video formats
  • Confidence scores for every word enable efficient human review
  • Entity formatting for numbers, dates, currencies automatically
  • Profanity and hesitation detection for compliance
  • On-device transcription in Adobe Premiere for professional work
  • No rate limits on Enterprise plans
  • Volume discounts automatically applied above 500 hours monthly
  • Startup Program offering up to $50k in credits
  • Multi-region cloud options for global deployment
  • Custom models and custom voice/language development for Enterprise

Speechmatics cons

  • Text-to-speech only supports English (British and American) as of 2026
  • Free tier capped at 2,400 minutes (40 hours) per month
  • Pro tier usage capped at 6,000 hours per month
  • Starting price $0.24 per hour may be high for small budgets
  • Custom dictionary content excluded from job configuration summaries
  • Standard model does not provide turnaround time benefits in real-time
  • Enhanced model costs more but Standard lacks speed advantages
  • TTS additional languages planned for 2025 but not yet released
  • Customer audio data never sent over network limits some cloud features
  • Additional discounts require 24,000 hours yearly usage minimum

Frequently asked questions about Speechmatics

What languages does Speechmatics support?

Speechmatics supports 56+ languages for transcription including Dutch, English, French, German, Irish, Italian, Portuguese, Spanish, Danish, Estonian, Finnish, Norwegian, Swedish, Belarusian, Bulgarian, Czech, Hungarian, Latvian, Lithuanian, Polish, Romanian, Russian, Slovakian, Slovenian, Ukrainian, Catalan, Galician, Greek, Maltese, Welsh, Esperanto, Interlingua, Arabic, Hebrew, Persian, Turkish, Uyghur, Bashkir, Bengali, Hindi, Marathi, Tamil, Urdu, Cantonese, Mandarin, Japanese, Korean, Mongolian, Malay, Indonesian, Thai, and Swahili. For AI translation, 69 language pairs are supported.

How much does Speechmatics cost?

Speechmatics pricing starts from $0.24 per hour of transcribed audio, falling below this at scale with Enterprise plans. The free tier includes 2,400 minutes (40 hours) per month. Pro tier includes volume discounts automatically applied above 500 hours monthly (20% discount) and additional discounts from 24,000 hours annually. Text-to-speech is $0.011 per 1,000 characters. Enterprise customers receive custom pricing with dedicated support.

Can Speechmatics transcribe phone numbers, postcodes, and account numbers accurately?

Yes. Speechmatics is purpose-built for alphanumeric accuracy, hitting 96.9% sequence accuracy on character strings, 98.0% on digits, and 85.4% on mixed alphanumerics. This means phone numbers, postcodes, account numbers, SKUs, and booking references land correctly the first time, which is critical for contact centers, voice agents, logistics, and workflows where misheard letters or digits cause callbacks or failed transactions.

What deployment options does Speechmatics offer?

Speechmatics offers four deployment options: SaaS (cloud), on-premises, on-device, and hybrid. Users can choose virtual appliance, Docker containers, or preconfigured virtual appliances. On-device deployment provides ultra-low latency and maximum data privacy, ideal for limited connectivity scenarios. On-premises meets architecture, security, and compliance needs by hosting the API in your own environment.

What is Standard vs Enhanced accuracy model?

Enhanced provides unbeatable accuracy and is best-in-class across all languages when accuracy is a must-have. Standard provides great accuracy when file turnaround time or cost-control are priorities. Note that the Standard model does not provide turnaround time benefits when using real-time speech-to-text. Different models can be used for different jobs, allowing you to choose the right model for each task.

Does Speechmatics support real-time transcription?

Yes, Speechmatics offers fast, reliable real-time speech-to-text in 56+ languages with sub-500ms latency. It is the fastest real-time speech-to-text engine, delivering final transcripts in under one second. The free tier includes 2 concurrent real-time sessions, while Pro tier includes 50 concurrent real-time sessions with no rate limits.

What is speaker diarization and does Speechmatics support it?

Speaker diarization accurately separates and tracks multiple speakers, even in overlapping, messy conversations. Speechmatics supports speaker labeling for each word, available for both batch and real-time transcription. This feature helps identify who said what and when, improving transcript comprehensibility for contact centers, meeting platforms, and media applications.

Is Speechmatics compliant with privacy and healthcare regulations?

Yes, Speechmatics is ISO/IEC 27001:2022 accredited, GDPR compliant, fully HIPAA compliant for healthcare use cases, and SOC 2 Type II-certified. Data is encrypted in transit and at rest with compliance-ready infrastructure. The platform does not log customer data as standard and never sends customer audio data over the network, only recording metadata about transcriber activity.

Can I customize Speechmatics for my industry-specific terminology?

Yes, Speechmatics allows customization through custom dictionary injection for up to 1,000 domain-specific terms for accurate recognition of names, jargon, acronyms, and branded terms. The platform also offers custom models, custom voice development, and custom language development for Enterprise customers. Finance language packs optimized to industry terminology are available now with more to follow.

How does volume discounting work at Speechmatics?

Volume discounts are automatically applied on billable usage above 500 hours for each type of Speech-To-Text in a given month. For example, 500 hours at base rate plus 300 hours at 20% discount for 800 hours total usage. Additional discounts are available starting from 24,000 hours usage per year. Different models can be used for different jobs, and discounts apply per model type. Pro tier is capped at 6,000 hours per month.

Categories

Use cases

Browse all AI tools on NeedAnAI