Voxtral TTS

Mistral's open-weight TTS that beat ElevenLabs in blind tests

Video & Audio Open Source Has API Open Source
Researched · Published · Reviewed
RECATOOLS Score
7.4 / 10
Capability
7
Value for money
9
Ease of use
6
ASEAN readiness
6
API quality
7
Founded
HQ
Users
Launched
Developer

Overview

Voxtral TTS is Mistral's 4.1B-parameter open-weights text-to-speech model (released March 2026) generating expressive, multilingual speech in nine languages and cloning voices from 2-3 seconds of audio. Weights are on Hugging Face under a non-commercial license; also served via Mistral's API.

Advertisement

Pricing

Pricing shown for reference only. These figures reflect RECATOOLS research as of 12 Jul 2026 and may be out of date or incomplete. This is not financial or purchasing advice — always confirm the current price on the provider’s official website before making any decision.

Free
Free
Free tier with core features.

What you can produce with Voxtral TTS

  • Multilingual TTS in 9 languages (incl. Hindi, Arabic)
  • Voice cloning from 2-3 seconds of reference audio
  • ~70ms model latency on an H200 GPU
  • Open weights on Hugging Face (CC BY-NC 4.0, non-commercial)
  • Served via Mistral API and Mistral Studio
  • Free tier for testing
Advertisement

ASEAN Perspective

Voxtral TTS in Southeast Asia

ASEAN-region availability and pricing notes coming soon. Drop the editorial team a note via /contact/ if you can supply local context (Singapore/Malaysia/Indonesia/Thailand/Vietnam).

RECATOOLS Verdict

Voxtral TTS's headline claim checks out better than most 'beats X' marketing: in Mistral's own human-preference testing, listeners picked Voxtral over ElevenLabs' flagship voices 58.3% of the time, rising to 68.4% in zero-shot voice-cloning comparisons, with roughly even results on steerability. Latency is real too -- about 70ms on an H200, competitive with ElevenLabs Flash. The catches: those benchmarks are self-reported and hadn't yet shown up on independent leaderboards like Artificial Analysis's Speech Arena as of these reports, and the open-weight license is CC BY-NC 4.0 -- non-commercial only, so shipping Voxtral in a paid product means going through the metered API instead ($16 per million characters) or negotiating a commercial license. No Southeast Asian languages in the initial nine. Strong pick for cost-conscious voice-agent builders who can live with API pricing or a non-commercial self-hosted deployment.

Independent AI-assisted assessment by RECATOOLS.

What people say

Mistral shipped Voxtral TTS in late March 2026 as a 4.1-billion-parameter model, positioned as the audio-output sibling to its existing Voxtral transcription (ASR) line -- filling out a full voice stack rather than starting from scratch.

The comparison everyone reaches for is ElevenLabs, and Mistral leaned into that directly. In blind human-preference tests using each model's flagship built-in voices, Voxtral won 58.3% of comparisons; in zero-shot voice cloning -- cloning a voice from a short reference clip rather than using a preset -- Voxtral's win rate rose to 68.4%. Against ElevenLabs v3 specifically, results were closer to even on explicit prompt-based steering, with a slight Voxtral edge on implicit (inferred-from-context) steering. Latency numbers are competitive rather than dominant: Mistral reports about 70ms of model latency on an H200 GPU versus roughly 75ms for ElevenLabs Flash v2.5, with end-to-end time-to-first-audio from the API around 0.8 seconds over PCM.

The important caveat, flagged by multiple independent writeups including DataCamp's: these are Mistral's own benchmarks. No third party had published a matching independent comparison as of these reports, and the Artificial Analysis Speech Arena leaderboard -- a widely-cited independent TTS ranking -- hadn't added Voxtral TTS yet at time of writing.

Community reaction moved fast on the tooling side: within weeks of release, developers had published a pure-C inference implementation (voxtral-tts.c, zero dependencies beyond libc), a Rust/Burn implementation for native and browser use, a Rust/vLLM-Omni server, and a desktop GUI wrapper -- the kind of ecosystem activity that tends to show up only around models people actually want to self-host.

Licensing is the practical sticking point. The open weights on Hugging Face are released under CC BY-NC 4.0 -- free to use and modify, but non-commercial only -- so a company wanting to ship Voxtral-powered audio in a paid product either pays for Mistral's hosted API ($0.016 per 1,000 characters, roughly $16 per million) or needs a separate commercial agreement. Self-hosting the open weights currently requires vLLM-Omni. Language coverage is nine languages including Hindi and Arabic; no Southeast Asian language is in the initial set.

Summary of public user & expert reviews, compiled by RECATOOLS.

About this listing

Researched on
Published on
Last reviewed

This entry was compiled from publicly available data including Voxtral TTS's official website, press releases, documentation, and reputable third-party publications. RECATOOLS is not affiliated with Voxtral TTS unless explicitly stated.

Data accuracy

Third-party AI tools update their pricing, features, availability, and policies frequently. Information here may be outdated by the time you read this — we make reasonable efforts to keep listings current, but cannot guarantee absolute accuracy.

For the latest details, please refer to Voxtral TTS directly →

Spotted something out of date? Suggest an update →

Advertisement