Hume AI

Empathic voice AI that responds to emotion

Video & Audio Freemium Has API
Researched · Published · Reviewed
RECATOOLS Score
7.2 / 10
Capability
8
Value for money
6
Ease of use
6
ASEAN readiness
5
API quality
8
Founded
2021
HQ
New York, USA

Overview

Hume AI builds empathic voice interfaces — language models that recognize and respond to emotional cues in speech. The EVI (Empathic Voice Interface) API lets developers add voice agents that adjust tone, pacing and content based on the user's emotional state. Research-driven, founded by ex-Google DeepMind researcher Alan Cowen.

Advertisement

Use cases

Empathic chatbots Mental-health support Customer-service voice agents UX research

What you can produce with Hume AI

  • Build a real-time voice agent with the EVI 3 API that detects frustration, hesitation or excitement in a caller's speech and adapts its tone and responses accordingly.
  • Generate expressive text-to-speech with Octave by directing performance in plain language, such as instructing the voice to sound sympathetic, sarcastic or urgent for a given line.
  • Design a custom brand voice from a text description, or create character voices that stay consistent across long-form narration.
  • Analyse recorded speech or video with the Expression Measurement API to quantify emotional signals for research, coaching or QA scoring of support calls.
  • Pair EVI's voice and empathy layer with your own LLM, so your existing model generates the words while Hume handles listening, emotion and speech.
  • Handle natural conversational behaviour — interruptions, back-channels and turn-taking — over a WebSocket connection without building your own audio pipeline.
Advertisement
RECATOOLS Verdict

Hume AI specialises in emotionally expressive voice, with its Empathic Voice Interface (EVI) and speech models that detect vocal emotion and respond with natural, prosody-rich speech. It is a genuine differentiator for voice agents, companions, coaching, and accessibility apps where tone and empathy matter, and the developer API is well documented for building real-time voice experiences. It suits developers building conversational voice products.

Caveats: emotion inference is scientifically contested and can be inaccurate or culturally biased, so it should be used carefully and not as ground truth, especially in high-stakes contexts. It is a builder platform, not an end-user app, and usage-based pricing scales with audio volume. English-strongest; multilingual and ASEAN-language support is more limited than text models. Strong API is its best asset.

Independent AI-assisted assessment by RECATOOLS.

What people say

Hume AI remains the research-flavoured contender in voice AI, built around founder Alan Cowen's work on emotion science at this New York lab. EVI, whose third generation shipped in May 2025, is a speech-to-speech model that detects emotional cues in a caller's voice and adjusts its own tone and content in response, and can speak in any of the 100,000-plus custom voices created on the platform without fine-tuning. Octave, now Octave 2, is a text-to-speech engine that interprets meaning — it can act out characters and shift register on instruction. Over 100,000 developers and businesses were on the APIs as of late 2025, across support, health, education, gaming, automotive and robotics.

Reviewers consistently say EVI picks up subtle shifts — stress, hesitation, excitement — that other voice stacks flatten, and users comparing it against ElevenLabs and PlayHT say the voices modulate in response to how the speaker sounds, not just what they say. In a blind study with 180 human raters Octave was preferred over ElevenLabs' TTS. Latency reads well for conversation: EVI responds in under 300 milliseconds and Octave 2 generates in under 200, and the Octave 2 launch halved per-character cost. SDK coverage across React, TypeScript, Python, Swift and .NET, and the ability to pair EVI with Claude, GPT, Gemini or Llama, earn their own praise.

This is an API-first platform with a genuine learning curve, far from plug-and-play for non-developers and small teams, with limited off-the-shelf integrations next to rivals. Costs are usage-based per minute — overage around $0.06-0.07 — cheap per call but quick to add up for always-on agents, and pricing draws criticism for inflexibility at small scale. Reviewers who tested the consumer-facing voice app praised the realism while concluding it is not quite there: empathic responses can feel uncanny or miss, because inferring emotion from prosody is inherently probabilistic. For ultra-low latency, ElevenLabs' Flash model at around 75ms still wins.

Hume fits engineering teams building voice agents where emotional register matters — support lines, coaching, health check-ins, game NPCs, character-driven media — and who are comfortable owning an API integration.

Summary of public user & expert reviews, compiled by RECATOOLS.

About this listing

Researched on
Published on
Last reviewed

This entry was compiled from publicly available data including Hume AI's official website, press releases, documentation, and reputable third-party publications. RECATOOLS is not affiliated with Hume AI unless explicitly stated.

Data accuracy

Third-party AI tools update their pricing, features, availability, and policies frequently. Information here may be outdated by the time you read this — we make reasonable efforts to keep listings current, but cannot guarantee absolute accuracy.

For the latest details, please refer to Hume AI directly →

Spotted something out of date? Suggest an update →

Advertisement