Voxtral TTS
Mistral's open-weight TTS that beat ElevenLabs in blind tests
Overview
Voxtral TTS is Mistral's 4.1B-parameter open-weights text-to-speech model (released March 2026) generating expressive, multilingual speech in nine languages and cloning voices from 2-3 seconds of audio. Weights are on Hugging Face under a non-commercial license; also served via Mistral's API.
Pricing
Pricing shown for reference only. These figures reflect RECATOOLS research as of 4 Sep 2026 and may be out of date or incomplete. This is not financial or purchasing advice — always confirm the current price on the provider’s official website before making any decision.
What you can produce with Voxtral TTS
- Multilingual TTS in 9 languages (incl. Hindi, Arabic)
- Voice cloning from 2-3 seconds of reference audio
- ~70ms model latency on an H200 GPU
- Open weights on Hugging Face (CC BY-NC 4.0, non-commercial)
- Served via Mistral API and Mistral Studio
- Free tier for testing
Voxtral TTS's headline claim checks out better than most 'beats X' marketing: in Mistral's own human-preference testing, listeners picked Voxtral over ElevenLabs' flagship voices 58.3% of the time, rising to 68.4% in zero-shot voice-cloning comparisons, with roughly even results on steerability. Latency is real too -- about 70ms on an H200, competitive with ElevenLabs Flash. The catches: those benchmarks are self-reported and hadn't yet shown up on independent leaderboards like Artificial Analysis's Speech Arena as of these reports, and the open-weight license is CC BY-NC 4.0 -- non-commercial only, so shipping Voxtral in a paid product means going through the metered API instead ($16 per million characters) or negotiating a commercial license. No Southeast Asian languages in the initial nine. Strong pick for cost-conscious voice-agent builders who can live with API pricing or a non-commercial self-hosted deployment.
What people say
Mistral shipped Voxtral TTS in late March 2026 as a 4.1-billion-parameter model, positioned as the audio-output sibling to its existing Voxtral transcription (ASR) line -- filling out a full voice stack rather than starting from scratch.
The comparison everyone reaches for is ElevenLabs, and Mistral leaned into that directly. In blind human-preference tests using each model's flagship built-in voices, Voxtral won 58.3% of comparisons; in zero-shot voice cloning -- cloning a voice from a short reference clip rather than using a preset -- Voxtral's win rate rose to 68.4%. Against ElevenLabs v3 specifically, results were closer to even on explicit prompt-based steering, with a slight Voxtral edge on implicit (inferred-from-context) steering. Latency numbers are competitive rather than dominant: Mistral reports about 70ms of model latency on an H200 GPU versus roughly 75ms for ElevenLabs Flash v2.5, with end-to-end time-to-first-audio from the API around 0.8 seconds over PCM.
The important caveat, flagged by multiple independent writeups including DataCamp's: these are Mistral's own benchmarks. No third party had published a matching independent comparison as of these reports, and the Artificial Analysis Speech Arena leaderboard -- a widely-cited independent TTS ranking -- hadn't added Voxtral TTS yet at time of writing.
Community reaction moved fast on the tooling side: within weeks of release, developers had published a pure-C inference implementation (voxtral-tts.c, zero dependencies beyond libc), a Rust/Burn implementation for native and browser use, a Rust/vLLM-Omni server, and a desktop GUI wrapper -- the kind of ecosystem activity that tends to show up only around models people actually want to self-host.
Licensing is the practical sticking point. The open weights on Hugging Face are released under CC BY-NC 4.0 -- free to use and modify, but non-commercial only -- so a company wanting to ship Voxtral-powered audio in a paid product either pays for Mistral's hosted API ($0.016 per 1,000 characters, roughly $16 per million) or needs a separate commercial agreement. Self-hosting the open weights currently requires vLLM-Omni. Language coverage is nine languages including Hindi and Arabic; no Southeast Asian language is in the initial set.
Summary of public user & expert reviews, compiled by RECATOOLS.
About this listing
This entry was compiled from publicly available data including Voxtral TTS's official website, press releases, documentation, and reputable third-party publications. RECATOOLS is not affiliated with Voxtral TTS unless explicitly stated.
Third-party AI tools update their pricing, features, availability, and policies frequently. Information here may be outdated by the time you read this — we make reasonable efforts to keep listings current, but cannot guarantee absolute accuracy.
For the latest details, please refer to Voxtral TTS directly →
Spotted something out of date? Suggest an update →
Alternatives to Voxtral TTS
More in Video & Audio