Amazon Polly
Amazon Polly is a text to speech software solution offered by Amazon Web Services. It converts text into lifelike speech, allowing users to...
Last verified:
What is Amazon Polly?
Amazon Polly is a fully managed text-to-speech service that converts text into an audio stream and generates voice on demand. It uses deep learning, neural networks, and generative voice engines to turn articles, web pages, PDF documents, and other text into lifelike speech.
It offers dozens of voices across many languages, with male and female options in many language variants. The site emphasizes natural-sounding, human-style voices that can be used to build speech-enabled applications for global audiences.
Polly is designed for developers and teams that want to add speech to apps, websites, RSS feeds, videos, mobile apps, IoT products, IVR systems, and accessibility tools. It supports speech generation, voice output customization, and use cases such as voiceovers, customer engagement, and multilingual dubbing.
The service also focuses on flexibility and control. Users can use SSML for pronunciation, emphasis, phrasing, and intonation, custom lexicons for specialized words, and standard audio formats like MP3 and OGG for storage, redistribution, analysis, or archiving.
Amazon Polly pricing
Pricing model: Free
Amazon Polly uses pay-as-you-go pricing based on the number of characters converted to speech or Speech Marks metadata. Standard voices cost $4.00 per 1 million characters outside the free tier, Neural voices cost $16.00 per 1 million characters, Long-Form voices cost $100.00 per 1 million characters, and Generative voices cost $30 per 1 million characters for speech requests. Caching and replaying generated speech costs nothing extra. The free tier includes 5 million characters per month for Standard voices, 1 million characters per month for Neural voices for the first 12 months, 500 thousand characters per month for Long-Form voices for the first 12 months, and 100 thousand characters per month for Generative voices for the first 12 months. New AWS customers can also receive up to $200 in AWS Free Tier credits starting July 15, 2025, with a free plan available for 6 months after account creation. AWS also notes GovCloud pricing separately, with Standard voices at $4.80 per 1 million characters and Neural voices at $19.20 per 1 million characters outside the free tier.
Amazon Polly pros
- Converts text into lifelike speech
- Fully managed AWS service
- Dozens of voices
- Many language options
- Male and female voice choices
- Supports SSML markup
- Custom lexicons for pronunciation
- Multiple voice engines
- Neural and generative voices
- Low-latency speech generation
- Can save output as MP3
- Can save output as OGG
- Caching and replay at no extra cost
- Useful for accessibility use cases
- Works for global applications
Amazon Polly cons
- Pricing is character-based
- Speech Marks are billed by character too
- Free tier varies by engine
- Free tier is time-limited for some voices
- SSML can cause errors if invalid
- Not all SSML tags work for all voices
- Higher-end voices cost much more
- Generative voices do not support Speech Marks pricing details
- Cloud service requires AWS account and integration
Frequently asked questions about Amazon Polly
What does Amazon Polly do?
Amazon Polly is a text-to-speech service that converts text into lifelike speech and outputs audio streams. It is built to help developers add voice capabilities to apps, websites, videos, accessibility tools, and other speech-enabled experiences.
What kinds of voices does Amazon Polly offer?
Amazon Polly offers dozens of lifelike voices across many languages and language variants, including male and female voices. The site says these voices are created using native speakers and are intended to help applications sound natural across different regions.
Can Amazon Polly use SSML?
Yes. Polly supports Speech Synthesis Markup Language, which lets you control pronunciation, emphasis, phrasing, intonation, pace, and style. The service also supports many SSML tags, but unsupported tags can produce errors, and tag availability can differ by voice type.
Can I customize pronunciation in Amazon Polly?
Yes. The site says Polly supports custom lexicons, which let you change how acronyms, company names, internal terms, and other words are pronounced. This is useful when the default pronunciation does not match your domain or brand language.
What audio formats can Amazon Polly produce?
The site says Polly can store and redistribute speech output in standard audio formats such as MP3 and OGG. It also supports streaming speech output and lets you cache generated speech for later reuse.
How does Amazon Polly pricing work?
Pricing is based on the number of characters processed for speech or Speech Marks metadata. Standard, Neural, Long-Form, and Generative voices each have different per-million-character rates, and cached playback of generated speech is free.
Is there a free tier for Amazon Polly?
Yes. The site says the free tier includes 5 million Standard-voice characters per month, 1 million Neural-voice characters per month for the first 12 months, 500 thousand Long-Form characters per month for the first 12 months, and 100 thousand Generative-voice characters per month for the first 12 months. AWS also mentions up to $200 in Free Tier credits for new customers starting July 15, 2025.
Who is Amazon Polly for?
Amazon Polly is aimed at developers, product teams, and organizations that want to add speech to software and content. The site highlights use cases such as websites, mobile apps, IoT devices, RSS feeds, videos, voice response systems, games, accessibility tools, and multilingual dubbing.
Does Amazon Polly retain submitted text content?
AWS states that content security, trust, and privacy are priorities and that Amazon Polly does not retain the content of text submissions. The service is positioned so generated speech can be stored locally or in standard audio files for reuse.
What makes Amazon Polly suitable for global applications?
Polly offers voices in many languages and language variants, with native-speaker voice models and multiple voice choices per language in many cases. This makes it useful for applications that need localized speech experiences across different countries and markets.