Conformer2

Conformer2 is Conformer-2 is an advanced automatic speech recognition AI model developed as a successor to Conformer-1. It's designed with robust improvements for decoding pr...

Last verified:

Visit Conformer2

What is Conformer2?

Conformer2 is Conformer-2 is AssemblyAI's latest AI model for automatic speech recognition (ASR), trained on 1.1 million hours of English audio data. It extends Conformer-1 with significant improvements in proper noun recognition, alphanumeric transcription accuracy, and robustness to noisy audio conditions. The model features 450 million parameters and is now the default speech recognition model on AssemblyAI's API, requiring no changes for existing users to benefit from the upgrade.

Conformer2 pricing

Pricing model: Free

AssemblyAI offers a free tier with $50 in free credits upon signup, equivalent to approximately 185 hours of pre-recorded audio transcription or 333 hours of streaming audio. Free accounts can start 5 new streams per minute. After using the free credits, users pay as-you-go with no commitments required. Pre-recorded Speech-to-Text pricing: Universal-2 at $0.15/hour and Universal-3 Pro at $0.21/hour. Streaming Speech-to-Text (Universal-Streaming) costs $0.15/hour with session-based billing. Volume discounts and custom tiered pricing are available by contacting sales for enterprise needs including custom rate limits, SLAs, and dedicated support.

Conformer2 pros

  • 31.7% improvement on alphanumeric transcription accuracy
  • 6.8% improvement on Proper Noun Error Rate
  • 12.0% improvement in robustness to noise
  • Up to 55% faster than Conformer-1 depending on audio duration
  • Trained on 1.1 million hours of English audio data
  • 450 million parameters for enhanced performance
  • Default speech recognition model on AssemblyAI API
  • No changes required for existing API users to upgrade
  • Model ensembling with multiple strong teacher models
  • Reduced variance in transcription results
  • 0% CER on files where Conformer-1 made mistakes
  • Transcription time for hour-long file down to 1.85 minutes
  • Strong performance on real-world audio conditions
  • Effective for call centers, podcasts, broadcasts, and webinars
  • Improved readability and consistency across transcripts

Conformer2 cons

  • Only supports English language
  • No significant improvements in Word Error Rate (WER)
  • Larger model size may increase compute requirements
  • Training speed ~1.6x faster only on in-house hardware
  • Asymptotic limit reached on bootstrapping teacher models
  • Diminishing returns observed on current research approach
  • Requires API token signup for free testing
  • Free tier limited to $50 in audio processing credits

Frequently asked questions about Conformer2

What is Conformer-2?

Conformer-2 is AssemblyAI's latest AI model for automatic speech recognition, trained on 1.1 million hours of English audio data. It improves on Conformer-1 with 31.7% better alphanumeric transcription, 6.8% better proper noun recognition, and 12.0% better noise robustness, while being up to 55% faster despite having 450 million parameters.

Is Conformer-2 the default model on AssemblyAI API?

Yes, Conformer-2 is already the default speech recognition model on AssemblyAI's API. Current API users will automatically be switched to Conformer-2 and start seeing better performance with no changes required on their end.

How much faster is Conformer-2 compared to Conformer-1?

Conformer-2 is faster than Conformer-1 by up to 55% depending on the duration of the audio file. For an hour-long file, transcription time dropped from 4.01 minutes to 1.85 minutes, a 53.7% latency reduction in the inference pipeline.

What languages does Conformer-2 support?

Conformer-2 is trained on English audio data and supports English language only. The blog explicitly states it is trained on 1.1M hours of English audio data.

How can I try Conformer-2?

The easiest way to try Conformer-2 is through AssemblyAI's Playground, where you can upload a file or enter a YouTube link to see a transcription in just a few clicks. You can also sign up for a free API token and use the Docs or welcome Colab to start using the API directly.

What is the speech_threshold parameter?

The speech_threshold is a new API parameter introduced with Conformer-2 that enables users to set a threshold for the proportion of speech that must be present in an audio file for it to be processed. The API automatically rejects audio files with speech proportions lesser than the set threshold, helping users control costs with files like sleep podcasts, instrumental music, and empty audio files.

What training techniques were used for Conformer-2?

Conformer-2 uses model ensembling with multiple strong teacher models instead of a single teacher, leveraging noisy student-teacher training (NST). This variance reduction technique creates a more robust model exposed to wider distribution of behaviors, where failure cases of individual models are subdued by successes of other ensemble models.

What is Proper Noun Error Rate (PPNER)?

PPNER is a novel metric crafted by AssemblyAI specifically to measure model performance for proper nouns. It uses a character-based metric called Jaro-Winkler similarity to quantify important errors like incorrectly transcribing someone's name or address, which are more consequential than generic word errors that WER doesn't differentiate.

How does Conformer-2 perform on alphanumeric data?

Conformer-2 shows a 30.7% relative reduction in mean Character Error Rate (CER) on alphanumeric datasets. It gets 0% CER on some files where Conformer-1 made mistakes and shows reduced variance, making it much more likely to avoid significant mistakes when transcribing credit card numbers, confirmation codes, and other numerical data.

What hardware was used to train Conformer-2?

Conformer-2 was trained on AssemblyAI's own GPU compute cluster of 80GB-A100s using a fault-tolerant Slurm scheduler for cluster management and job scheduling. This in-house hardware made training speed approximately 1.6x faster than comparable cloud provider infrastructure and gives flexibility for constant experimentation.

Categories

Use cases

Browse all AI tools on NeedAnAI