Voxtral TTS
Mistral's open-weight TTS that beat ElevenLabs in blind tests
Overview
Voxtral TTS is Mistral's 4.1B-parameter open-weights text-to-speech model (released March 2026) generating expressive, multilingual speech in nine languages and cloning voices from 2-3 seconds of audio. Weights are on Hugging Face under a non-commercial license; also served via Mistral's API.
Pricing
Pricing shown for reference only. These figures reflect RECATOOLS research as of 12 Jul 2026 and may be out of date or incomplete. This is not financial or purchasing advice — always confirm the current price on the provider’s official website before making any decision.
What you can produce with Voxtral TTS
- Multilingual TTS in 9 languages (incl. Hindi, Arabic)
- Voice cloning from 2-3 seconds of reference audio
- ~70ms model latency on an H200 GPU
- Open weights on Hugging Face (CC BY-NC 4.0, non-commercial)
- Served via Mistral API and Mistral Studio
- Free tier for testing
ASEAN Perspective
Voxtral TTS in Southeast Asia
ASEAN-region availability and pricing notes coming soon. Drop the editorial team a note via /contact/ if you can supply local context (Singapore/Malaysia/Indonesia/Thailand/Vietnam).
Voxtral TTS's headline claim checks out better than most 'beats X' marketing: in Mistral's own human-preference testing, listeners picked Voxtral over ElevenLabs' flagship voices 58.3% of the time, rising to 68.4% in zero-shot voice-cloning comparisons, with roughly even results on steerability. Latency is real too -- about 70ms on an H200, competitive with ElevenLabs Flash. The catches: those benchmarks are self-reported and hadn't yet shown up on independent leaderboards like Artificial Analysis's Speech Arena as of these reports, and the open-weight license is CC BY-NC 4.0 -- non-commercial only, so shipping Voxtral in a paid product means going through the metered API instead ($16 per million characters) or negotiating a commercial license. No Southeast Asian languages in the initial nine. Strong pick for cost-conscious voice-agent builders who can live with API pricing or a non-commercial self-hosted deployment.
What people say
Mistral shipped Voxtral TTS in late March 2026 as a 4.1-billion-parameter model, positioned as the audio-output sibling to its existing Voxtral transcription (ASR) line -- filling out a full voice stack rather than starting from scratch.
The comparison everyone reaches for is ElevenLabs, and Mistral leaned into that directly. In blind human-preference tests using each model's flagship built-in voices, Voxtral won 58.3% of comparisons; in zero-shot voice cloning -- cloning a voice from a short reference clip rather than using a preset -- Voxtral's win rate rose to 68.4%. Against ElevenLabs v3 specifically, results were closer to even on explicit prompt-based steering, with a slight Voxtral edge on implicit (inferred-from-context) steering. Latency numbers are competitive rather than dominant: Mistral reports about 70ms of model latency on an H200 GPU versus roughly 75ms for ElevenLabs Flash v2.5, with end-to-end time-to-first-audio from the API around 0.8 seconds over PCM.
The important caveat, flagged by multiple independent writeups including DataCamp's: these are Mistral's own benchmarks. No third party had published a matching independent comparison as of these reports, and the Artificial Analysis Speech Arena leaderboard -- a widely-cited independent TTS ranking -- hadn't added Voxtral TTS yet at time of writing.
Community reaction moved fast on the tooling side: within weeks of release, developers had published a pure-C inference implementation (voxtral-tts.c, zero dependencies beyond libc), a Rust/Burn implementation for native and browser use, a Rust/vLLM-Omni server, and a desktop GUI wrapper -- the kind of ecosystem activity that tends to show up only around models people actually want to self-host.
Licensing is the practical sticking point. The open weights on Hugging Face are released under CC BY-NC 4.0 -- free to use and modify, but non-commercial only -- so a company wanting to ship Voxtral-powered audio in a paid product either pays for Mistral's hosted API ($0.016 per 1,000 characters, roughly $16 per million) or needs a separate commercial agreement. Self-hosting the open weights currently requires vLLM-Omni. Language coverage is nine languages including Hindi and Arabic; no Southeast Asian language is in the initial set.
Summary of public user & expert reviews, compiled by RECATOOLS.
About this listing
This entry was compiled from publicly available data including Voxtral TTS's official website, press releases, documentation, and reputable third-party publications. RECATOOLS is not affiliated with Voxtral TTS unless explicitly stated.
Third-party AI tools update their pricing, features, availability, and policies frequently. Information here may be outdated by the time you read this — we make reasonable efforts to keep listings current, but cannot guarantee absolute accuracy.
For the latest details, please refer to Voxtral TTS directly →
Spotted something out of date? Suggest an update →
Alternatives to Voxtral TTS
More in Video & Audio