MiniMax Audio
Studio-grade multilingual TTS and voice cloning from China's most-funded AI lab — 50+ languages, 10-second clone.
Overview
MiniMax Audio is the speech and music generation platform from MiniMax, a Shanghai-based AI lab founded in 2021 by former SenseTime researchers. It offers text-to-speech synthesis across 50+ languages and accents, voice cloning from as little as 10 seconds of reference audio, a voice isolator for noise removal, and an official library of 300+ customisable voices. The underlying Speech model family — currently at version 2.8 — secured the top position on the public TTS Arena leaderboard in 2025 and powers integrations in LiveKit, Vapi, and Pipecat. The platform is accessible at minimax.io/audio (formerly hailuo.ai) via a web interface and a REST API.
MiniMax Music 2.5, launched January 28 2026, extended the platform into AI music generation with paragraph-level structural control via 14 section tags (Intro, Verse, Chorus, Bridge, and more), studio-grade vocal synthesis featuring natural vibrato and chest-to-head resonance transitions, adaptive mixing across 100+ instruments, and full-track generation up to five minutes. Music 2.6 followed in April 2026, adding improved bass rendering and cover-style generation. Together, the audio and music products position MiniMax as one of the most capable multilingual AI voice platforms available via a public API.
Pricing
Pricing shown for reference only. These figures reflect RECATOOLS research as of 11 Jul 2026 and may be out of date or incomplete. This is not financial or purchasing advice — always confirm the current price on the provider’s official website before making any decision.
- 100K points/mo
- T2A v2 & v2-large API
- Voice cloning, 300+ voices
- Commercial license
- 300K points/mo
- Faster generation speed
- Voice cloning, 300+ voices
- Commercial license
- 1.1M points/mo
- Higher rate limits
- Voice cloning, 300+ voices
- Commercial license
- 3.3M points/mo
- Higher rate limits
- Priority support
- Commercial license
- 20M points/mo
- 800 voice slots
- Dedicated support
- Commercial license
- Unlimited RPM/TPM
- Priority access
- Custom voice slots
- Enterprise SLA
Use cases
What you can produce with MiniMax Audio
- Natural-sounding speech audio in 50+ languages generated from a plain-text input
- Cloned voice profile created from a 10-second reference recording, usable via API
- Full AI-produced music track up to 5 minutes with vocals, instrumentation, and mixing from tagged lyrics
- Low-latency sub-250ms streaming speech output for real-time voice agent integration
- Noise-cleaned vocal-isolated audio output via the Voice Isolator tool
- Batch long-form synthesis of up to 10 million characters per request for audiobook pipelines
ASEAN Perspective
MiniMax Audio in Southeast Asia
MiniMax Audio supports Vietnamese and Thai among its 50+ languages — two ASEAN languages with complex tonal structures where many Western TTS platforms struggle — and the MiniMax-Speech technical report specifically highlights these as benchmarked languages. Indonesian is also in the supported set, covering the region's two largest internet markets. However, dedicated ASEAN-language voice libraries are thinner than those for English, Mandarin, and Japanese, and sub-regional languages like Tagalog and Burmese are not prominently documented. For ASEAN developers building multilingual pipelines, MiniMax Audio offers a cost-competitive API with solid foundational support, but production quality for less-resourced regional languages should be validated with real samples before deployment.
MiniMax Audio is one of the strongest multilingual TTS platforms available as of mid-2026, particularly for East and Southeast Asian languages where Western competitors tend to underperform. The Speech model family topped the TTS Arena leaderboard in 2025, voice cloning works from just 10 seconds of audio, and the API pricing (from roughly $0.04 per 1,000 characters) undercuts ElevenLabs significantly at scale. MiniMax Music 2.5 adds a genuinely capable music generation layer with paragraph-level structural control and natural vocal synthesis that set it apart from earlier generative music tools.
Caveats are real, however. The English voice variety and emotional nuance still lag behind ElevenLabs in side-by-side tests. API documentation is less polished than established Western peers, and the platform's Chinese-first heritage means English support content and community resources are thinner. Music generation lacks stem export and in-session editing, and the free Music tier is effectively trial-only. Businesses requiring guaranteed commercial licensing should confirm terms directly, as licensing language is not prominently displayed. Copyright litigation filed by Disney, Universal, and Warner Bros. Discovery in 2025 is an open risk factor worth monitoring for enterprise buyers.
What people say
Speech-02 topped the Hugging Face TTS Arena in 2025; by mid-2026 it has slipped to roughly seventh or eighth place as Vocu V3.0, Sonic-series, and Gemini Flash TTS models pushed ahead. That drop matters less than it sounds — reviewers still single out MiniMax for East and Southeast Asian language naturalness, an area where Western TTS vendors routinely stumble.
Six subscription tiers run from $5/month (100K audio points) up to a $999/month Business plan with 20M points and 800 voice slots, plus custom enterprise pricing above that. API rates land around $60–100 per million characters, and voice cloning from a 10-second sample costs about $1.50 per clone — both comfortably undercut ElevenLabs at volume. Music 2.5 and 2.6 extended the platform into full-track generation with section-tag structural control, though there's still no stem export or in-session editing, and the free Music tier is trial-only.
English voice variety and API documentation lag the established Western players, and support content skews toward a Chinese-first audience. The bigger overhang for enterprise buyers: Disney, Universal, and Warner Bros. Discovery are suing MiniMax over its Hailuo video/image product, not the audio tools specifically, but the litigation (MiniMax lost its motion to dismiss in May 2026) is company-level risk worth tracking before signing a contract.
Summary of public user & expert reviews, compiled by RECATOOLS.
Notable facts
- MiniMax listed on the Hong Kong Stock Exchange on January 9 2026, surging 43% on debut to reach a valuation of $9.3 billion — one of the largest Chinese AI IPOs in recent years.
- The MiniMax Speech model is an Autoregressive Transformer that topped the public TTS Arena leaderboard in 2025, outperforming ElevenLabs Multilingual v2 on cross-lingual synthesis and tonal language accuracy.
- MiniMax Music 2.5 understands 14 structural section tags, giving users director-level control over song architecture before a single note is generated.
- The company name is a direct reference to the minimax decision algorithm from game theory, and all three founders came from SenseTime, one of China's largest computer vision firms.
Frequently asked questions
About this listing
This entry was compiled from publicly available data including MiniMax Audio's official website, press releases, documentation, and reputable third-party publications. RECATOOLS is not affiliated with MiniMax Audio unless explicitly stated.
Third-party AI tools update their pricing, features, availability, and policies frequently. Information here may be outdated by the time you read this — we make reasonable efforts to keep listings current, but cannot guarantee absolute accuracy.
For the latest details, please refer to MiniMax Audio directly →
Spotted something out of date? Suggest an update →
More in Video & Audio