Deepgram
Real-time speech-to-text platform
Overview
Deepgram sells developer APIs for real-time and batch speech-to-text (Nova-3, from $0.0048/min), text-to-speech (Aura) and voice agents; its Flux model builds sub-300ms end-of-turn detection directly into the transcription layer. Customers include Spotify, NASA and Citi.
Pricing
Pricing shown for reference only. These figures reflect RECATOOLS research as of 11 Jul 2026 and may be out of date or incomplete. This is not financial or purchasing advice — always confirm the current price on the provider’s official website before making any decision.
- All public models incl. Nova-3 and Flux
- Speech-to-text from $0.0048/min
- Community and Discord support
- Standard uptime SLAs
- Credits redeemed against actual usage
- Higher concurrency limits across APIs
- All public model endpoints
- Custom support and SLAs
- HIPAA and SOC 2 options
- Dedicated account management
Use cases
What you can produce with Deepgram
- Build a real-time voice agent with model-integrated end-of-turn detection using Deepgram's Flux model, eliminating the cut-off and awkward-pause problems common in voice AI pipelines.
- Transcribe call-center recordings at scale via the Nova-3 pre-recorded API, with automatic speaker diarization to identify who said what across multi-party calls.
- Generate natural-sounding synthetic voice responses using Aura-2 text-to-speech (sub-200ms time-to-first-byte) for IVR systems, audiobook narration, or accessibility read-aloud features.
- Stream live captions with sub-300ms transcription latency for webinars, live broadcasts, or accessibility tools using Deepgram's WebSocket streaming API.
- Process multilingual conversations — detecting and switching between 10 languages mid-call — using Flux Multilingual (released April 2026) without separate per-language model integrations.
- Extract structured insights from audio using Deepgram's Audio Intelligence features: automatic summarisation, topic detection, and sentiment analysis layered on top of transcriptions.
- Self-host Deepgram models on-premises or in a private cloud for enterprises with data-residency or air-gapped compliance requirements, using the same API interface as the cloud offering.
ASEAN Perspective
Deepgram in Southeast Asia
ASEAN-region availability and pricing notes coming soon. Drop the editorial team a note via /contact/ if you can supply local context (Singapore/Malaysia/Indonesia/Thailand/Vietnam).
Deepgram is the specialist pick for latency-sensitive voice applications. Nova-3 streams transcription from $0.0048/min with sub-300ms latency, and the Flux model (multilingual since April 2026) builds end-of-turn detection into the model itself, shaving 200–600ms off voice-agent response loops versus a separate VAD. Docs, SDKs and the API Playground are consistently praised, and billing granularity means short conversational clips often cost 30–40% less than rivals with similar headline rates. Caveats: it's developer infrastructure, not an end-user app; language coverage is uneven, with Chinese and low-resource Southeast Asian languages trailing English; add-ons like diarization bill separately; and the $200 credit is one-time, not a monthly free tier. The Growth plan's $4,000+ annual prepay is steep for early-stage teams. Globally available, ASEAN included.
What people say
A common newcomer complaint sets the tone: the $200 sign-up credit is a one-time allocation, not a recurring monthly free tier, and developers regularly discover this only when the balance runs dry. Past that speed bump, Deepgram's reputation among developers is about as good as it gets in speech infrastructure.
Nova-3 streams transcription from $0.0048/min at latency competitors struggle to match, and because Deepgram bills without rounding penalties, short conversational clips often come out 30–40% cheaper than rivals with similar headline rates. The newer Flux model bakes end-of-turn detection into the transcription layer itself — median under 300ms — saving 200–600ms per agent turn versus running a separate VAD, which is why Nova-3 plus Flux became the default voice-agent stack this year. Flux Multilingual, released April 2026, handles 10 languages with mid-call switching. Documentation, SDKs and the API Playground draw consistent praise, and enterprise teams migrating off incumbent cloud speech APIs report cost and accuracy gains, especially on noisy call-centre audio.
The criticisms are specific. Language coverage is uneven — Chinese and less-common languages trail English, Spanish and Portuguese noticeably, which matters for Southeast Asian deployments. Add-ons bill separately: diarization and redaction at $0.0020/min each, keyterm prompting at $0.0013/min, so real per-minute costs run above the sticker. The Growth plan wants a $4,000+ annual prepay, steep for early-stage startups, and occasional API-stability wobbles surface in developer forums. It's also pure infrastructure — non-developers wanting drag-and-drop transcription should look elsewhere.
For latency-sensitive voice products, though, this is the specialist choice: 200,000+ developers were on the platform as of the company's January 2025 count, and the 2026 releases have only widened its lead on turn-taking.
Summary of public user & expert reviews, compiled by RECATOOLS.
About this listing
This entry was compiled from publicly available data including Deepgram's official website, press releases, documentation, and reputable third-party publications. RECATOOLS is not affiliated with Deepgram unless explicitly stated.
Third-party AI tools update their pricing, features, availability, and policies frequently. Information here may be outdated by the time you read this — we make reasonable efforts to keep listings current, but cannot guarantee absolute accuracy.
For the latest details, please refer to Deepgram directly →
Spotted something out of date? Suggest an update →
Alternatives to Deepgram
More in Video & Audio