Hume AI
Hume AI is a unique AI suite designed to measure, comprehend, and enhance the influence of technology on human emotions. The platform featu...
Last verified:
What is Hume AI?
Hume AI is an empathic AI research lab and technology company that builds speech-language models with emotional intelligence to create technology that truly understands humanity. The platform provides open-source models, datasets, and evaluation APIs to embed emotional intelligence into voice models. Hume develops two primary APIs: the Empathic Voice Interface (EVI) for real-time voice interaction and Text-to-Speech (Octave TTS) for expressive speech synthesis. Their decades of research in multimodal emotional intelligence span 50+ languages, 48+ emotions, and 600+ voice descriptors.
EVI is Hume's flagship voice AI that understands and responds to human emotions in real-time, combining speech recognition, emotion detection, and natural language processing. It features interruptibility, pause/resume responses, chat resume with full context, chat history with timestamps and emotion data, dynamic variables, context injection for RAG, tool use for API connections, and audio reconstruction. EVI works seamlessly with any LLM including Claude, GPT, Gemini, Grok, Kimi K2, and Llama. The platform offers a curated library of 100+ expressive voices, voice cloning from samples, and voice design from natural language descriptions.
Octave TTS is the first text-to-speech system built on LLM intelligence that understands what words mean in context, unlocking expressiveness and nuance beyond conventional TTS. The platform is enterprise-ready with SOC 2 Type II security and HIPAA compliance, supporting custom SLAs, dedicated support, and volume pricing for business applications.
Hume AI is designed for developers building voice-first AI experiences, companies creating digital companions for seniors/kids/mental wellness, coaching and interviewing applications, digital assistants, creative tools for video/podcasting/audiobooks, education platforms, and digital avatars for apps/games/virtual experiences. Full SDK support is available for React, TypeScript, Python, .NET, Swift, and more.
Hume AI pricing
Pricing model: Freemium
Hume AI offers 7 pricing tiers: Free ($0/month) includes 10,000 TTS characters (~10 minutes), 5 EVI minutes, 1 concurrent connection, voice cloning (create only), 15 RPM rate limit, and Discord support. Starter ($3/month) includes 30,000 TTS characters (~30 minutes), 40 EVI minutes ($0.07/min), 5 concurrent connections, 20 projects, and voice cloning (create only). Creator ($14/month) includes 140,000 TTS characters (~140 minutes), 200 EVI minutes ($0.07/min), unlimited voice cloning, and commercial license. Pro ($70/month) includes 1M TTS characters, 1,200 EVI minutes, external LLMs support, and 10 concurrent connections. Scale ($200/month) includes 3.3M TTS characters, 5,000 EVI minutes, organization features, and 3 team seats. Business ($500/month) includes 10M TTS characters, 12,500 EVI minutes, and 5 team seats. Enterprise offers custom pricing with custom SLAs, dedicated support, and volume pricing. New accounts start on the free tier with $20 in credits. Overage rates: $0.15/1,000 TTS characters on Creator, $0.06/min on Pro, decreasing to $0.04/min on Business. Expression Measurement is billed pay-as-you-go (credits) per minute for audio/video, per image, and per word for text.
Hume AI pros
- Understands and responds to human emotions in real-time
- Works seamlessly with any LLM (Claude, GPT, Gemini, Grok, Llama, etc.)
- 300ms time to first byte for fast real-time responses
- Users can interrupt at any time like real conversations
- Programmatically pause and resume responses at any moment
- Pick up conversations where they left off with full context
- Complete transcripts with timestamps and emotion data
- Inject real-time data like user names and account info
- Add context mid-conversation for RAG and knowledge bases
- Connect to your APIs for appointments, orders, and more
- Reconstruct complete audio from any conversation on demand
- 100+ expressive voices in the Voice Library
- Clone any voice from a sample with user consent
- Design voices from natural language descriptions
- Full SDK support for React, TypeScript, Python, .NET, Swift
- SOC 2 Type II enterprise-grade security
- HIPAA compliant for healthcare applications
- Supports 50+ languages and 48+ emotions
- Streaming audio output starts playback in milliseconds
- No-code playground for testing without coding
Hume AI cons
- Free tier has very limited usage (5 EVI minutes, 10K TTS characters)
- Only 1 concurrent connection on free tier
- Overage rates can be expensive ($0.06-$0.15 per 1,000 TTS characters)
- EVI minutes limited on lower tiers (40 min on Starter at $3/month)
- External LLM usage adds supplemental billing costs
- Voice cloning on free/Starter is create-only with no commercial use
- 15 RPM rate limit on free tier
- Only Discord support on free and Starter plans
- Downgrades take effect at next renewal only
- Payment failures may pause API access until settled
Frequently asked questions about Hume AI
What is EVI (Empathic Voice Interface)?
EVI is Hume's voice AI that understands and responds to human emotions in real-time. It combines speech recognition, emotion detection, and natural language processing to create more natural, empathic conversations. EVI measures users' nuanced vocal modulations and responds using a speech-language model trained on millions of human interactions, uniting language modeling and text-to-speech with better EQ, prosody, end-of-turn detection, interruptibility, and alignment.
What is Octave TTS?
Octave TTS is the first text-to-speech system built on LLM intelligence. Unlike conventional TTS that merely reads words, Octave is a speech-language model that understands what words mean in context, unlocking a new level of expressiveness and nuance. All voices in Hume's platform are powered by Octave, which enables expressive, context-aware speech generation from both text and natural language descriptions.
Can I use my own LLM with EVI?
Yes, EVI works seamlessly with any language model including Claude, GPT, Gemini, Grok, Kimi K2, Llama, and more. You can bring your own LLM and switch models without changing your integration. If you use your own LLM API key or a custom language model, Hume does not charge for that LLM usage. However, using Hume managed external LLMs adds supplemental usage to your bill.
What voice options are available?
Hume offers three voice options: Voice Library with over 100 expressive voices designed by Hume available for immediate use, Voice Design to create custom voices using descriptive prompts and Octave's expressive generation, and Voice Cloning to clone a voice from a recorded or uploaded speech sample with user consent. Custom voices are stored in your account and only accessible to you. Voices can be used across both EVI and TTS.
What SDKs does Hume support?
Hume provides full SDK support for React, TypeScript, Python, .NET, and Swift. The React SDK integrates EVI into React apps with tools for audio recording, playback, and API interaction. The TypeScript SDK integrates Hume APIs into Node applications or frontend Web applications. The Python SDK offers async/sync clients, error handling, and streaming tools. The Swift SDK builds iOS and macOS apps with EVI voice chat, microphone capture, realtime playback, and TTS file streaming. The .NET SDK provides typed TTS clients, automatic retries, pagination, and configurable timeouts.
Is Hume AI enterprise-ready?
Yes, Hume AI is enterprise-ready with SOC 2 Type II enterprise-grade security with industry-leading practices, HIPAA compliance for building healthcare applications with confidence, and Enterprise plans with custom SLAs, dedicated support, and volume pricing. The platform is built for business deploying with confidence knowing EVI meets security, compliance, and scale requirements of enterprise applications.
What are the main use cases for EVI?
EVI has several primary use cases: Interviewing & Coaching to simulate lifelike interviews or leadership coaching sessions with dynamic tone adjustment, Digital Companions to build emotionally aware companions for seniors, kids, or mental wellness support, and Digital Assistants to respond with empathy and modulate tone to reduce user frustration or improve engagement. Additional use cases include creative tools for narration, education/coaching with emotionally varied voice, and digital avatars for AI-powered characters in apps, games, or virtual experiences.
How does billing work for Hume AI?
Hume uses two billing models: Subscription for TTS, EVI, and Voice features (one subscription includes all three with included usage and overage rates after exceeding limits), and Pay as you go (credits) for Expression Measurement billed usage by usage. New accounts start on the free tier with $20 in credits. Billing cycle starts the day you begin a paid subscription and renews monthly. Included usage and plan limits reset at the start of each cycle. Upgrades take effect immediately; downgrades take effect at next renewal. Overage billing applies when you exceed included usage.
What is the free tier included usage?
The Free tier includes 10,000 TTS characters per month (~10 minutes), 5 EVI minutes per month, 1 concurrent connection, voice cloning (create only, no commercial use), 15 RPM rate limit, and Discord support. New accounts start on the free tier with $20 in credits. This tier is designed for testing and learning but has very limited usage for production applications.
How do I get my API keys?
Each Hume account is provisioned with an API key and Secret key accessible from the Hume Portal. To get your keys: 1) Sign in to the Hume Portal and log in or create an account, 2) Navigate to the API keys page to view your keys. Hume APIs support two authentication strategies: API key strategy for server-side requests using the X-Hume-Api-Key header, and Token strategy for client-side requests where you obtain a temporary access token (expires after 30 minutes) first. API keys can be regenerated by clicking the Regenerate keys button, which permanently invalidates the current keys.