Whisper API
Whisper API is an AI-powered transcription tool that allows users to send audio files via an API and receive back accurate transcriptions using OpenAI's Whisper model.. [Paid]
Last verified:
What is Whisper API?
Whisper API is an audio and video transcription service built around OpenAI Whisper and Whisper v3, aimed at turning recordings into text quickly and accurately. It supports uploads through both an API and a simple web transcription tool, making it usable for developers and non-developers alike.
The service emphasizes fast turnaround, saying transcripts can be delivered in minutes and in some cases as little as one-tenth of the playback time. It supports over 100 languages, speaker detection / diarization, translations, summaries, timestamps, and multiple export formats.
Its main product is an OpenAI-compatible transcription API that can be integrated into applications with just a few lines of code. The website also highlights a free transcription tool for drag-and-drop uploads, plus a separate product called Transcripo for users who want a more advanced no-code transcription workflow.
Whisper API is positioned for podcast creators, meeting teams, developers, content teams, and anyone who needs speech turned into searchable text, subtitles, or downloadable transcript files. It is also marketed as a cost-conscious option for scaling transcription workloads.
Whisper API pricing
Pricing model: Freemium
The website says the service includes a free first month with 30 hours of transcription, and the free web tool has no credit card, no signup, and no watermark. After the trial, the homepage advertises pricing of $0.17 per hour. The site does not show a fuller tier table on the pages reviewed, but it positions Whisper API as an affordable usage-based transcription service.
Whisper API pros
- 30 hours free in the first month
- No credit card required for free trial
- No signup required for the free transcription tool
- OpenAI-compatible API
- Simple integration in minutes
- Supports audio and video files
- Supports 100+ languages
- Speaker diarization support
- Translation support
- Summaries from transcript content
- Timestamps support
- Exports text files
- Exports PDF transcripts
- Exports SRT subtitles
- Exports VTT subtitles
- Works with secure cloud processing
- Fast transcription turnaround
- Scales to large user volumes
- Low per-hour pricing after trial
- Web tool available for non-developers
Whisper API cons
- Free web tool limited to 20 MB files
- Free web tool is more limited than the API
- Some advanced features are pushed to Transcripo
- Pricing details beyond the headline rate are sparse
- Website does not show a broad tier breakdown
- No explicit enterprise plan details on the homepage
- No visible refund policy on the homepage
- No clear mention of API rate limits on the homepage
- No on-page accuracy guarantees or SLAs
- No detailed privacy or retention terms on the homepage
Frequently asked questions about Whisper API
What does Whisper API do?
Whisper API transcribes audio and video into text using OpenAI Whisper / Whisper v3. It can handle recordings such as podcasts, videos, meetings, interviews, and similar files, and it also supports subtitles, timestamps, speaker labels, translations, and transcript exports.
How much free usage do I get?
The homepage says you get the first month free with 30 hours of transcription. The free transcription page also says you can use the web tool without a credit card, without signup, and without a watermark.
What file types can I upload?
The transcription page says all audio and video file formats are supported, and it explicitly mentions common files like mp3 and wav. It also says the free upload tool should not exceed 20 MB.
Does it support multiple languages?
Yes. The website says it supports 100+ languages on the main homepage and 96+ languages on the free transcription page. That makes it suitable for multilingual speech-to-text workflows and translation use cases.
Can I get speaker labels and timestamps?
Yes. The site says the service supports speaker diarization, which labels multiple speakers in an audio file, and it also supports timestamps. Those features are highlighted as part of the broader transcription workflow.
Can I export subtitles or text files?
Yes. The transcription tool page says you can copy or download transcripts as text, PDF, or SRT / VTT for video subtitles. That makes the output usable for documentation, publishing, and video workflows.
Is there an API for developers?
Yes. Whisper API offers an OpenAI-compatible transcription API that can be integrated into applications. The site says developers can connect it quickly with just a few lines of code and use documentation and code examples to get started.
Who is Whisper API for?
It is aimed at developers building transcription features, as well as non-developers who just want to upload a file and get a transcript. The website specifically mentions podcasts, videos, meetings, interviews, and general speech-to-text use cases.
How fast is transcription?
The site says transcription can finish in minutes and, on the free tool page, claims it can process files in as little as one-tenth of playback time. It gives an example that a 10-minute file can finish in a few seconds.
What happens after the free trial?
The homepage says pricing becomes $0.17 per hour after the free month and 30 free transcription hours. The site presents this as a simple affordable pricing model rather than a large tiered plan structure.