MiniMax Audio

Studio-grade multilingual TTS and voice cloning from China's most-funded AI lab — 50+ languages, 10-second clone.

Video & Audio Freemium Has API
Researched · Published · Reviewed
RECATOOLS Score
8.1 / 10
Capability
8.8
Value for money
8.5
Ease of use
7.5
ASEAN readiness
7.8
API quality
7.5
Founded
2021
HQ
Shanghai, China
Users
Used across MiniMax's Hailuo AI and platform apps worldwide
Launched
Audio consolidated under minimax.io Nov 2024
Developer
MiniMax Group (listed on HKEX, January 2026)

Overview

MiniMax Audio is the speech and music generation platform from MiniMax, a Shanghai-based AI lab founded in 2021 by former SenseTime researchers. It offers text-to-speech synthesis across 50+ languages and accents, voice cloning from as little as 10 seconds of reference audio, a voice isolator for noise removal, and an official library of 300+ customisable voices. The underlying Speech model family — currently at version 2.8 — secured the top position on the public TTS Arena leaderboard in 2025 and powers integrations in LiveKit, Vapi, and Pipecat. The platform is accessible at minimax.io/audio (formerly hailuo.ai) via a web interface and a REST API.

MiniMax Music 2.5, launched January 28 2026, extended the platform into AI music generation with paragraph-level structural control via 14 section tags (Intro, Verse, Chorus, Bridge, and more), studio-grade vocal synthesis featuring natural vibrato and chest-to-head resonance transitions, adaptive mixing across 100+ instruments, and full-track generation up to five minutes. Music 2.6 followed in April 2026, adding improved bass rendering and cover-style generation. Together, the audio and music products position MiniMax as one of the most capable multilingual AI voice platforms available via a public API.

Advertisement

Pricing

Pricing shown for reference only. These figures reflect RECATOOLS research as of 11 Jul 2026 and may be out of date or incomplete. This is not financial or purchasing advice — always confirm the current price on the provider’s official website before making any decision.

Starter
$5/mo
100K audio points, entry tier
  • 100K points/mo
  • T2A v2 & v2-large API
  • Voice cloning, 300+ voices
  • Commercial license
Standard
$30/mo
300K audio points
  • 300K points/mo
  • Faster generation speed
  • Voice cloning, 300+ voices
  • Commercial license
Scale
$249/mo
3.3M audio points
  • 3.3M points/mo
  • Higher rate limits
  • Priority support
  • Commercial license
Business
$999/mo
20M audio points, 800 voice slots
  • 20M points/mo
  • 800 voice slots
  • Dedicated support
  • Commercial license
Custom
Custom
Enterprise scale & SLA
  • Unlimited RPM/TPM
  • Priority access
  • Custom voice slots
  • Enterprise SLA

Use cases

Multilingual video voiceovers and dubbing for ASEAN content creators Developer API integration for voice agents, chatbots, and smart device assistants E-learning narration across English, Chinese, Japanese, Korean, Indonesian, Vietnamese, and Thai AI music generation for podcast intros, social-media content, and royalty-free background tracks Voice cloning for personalised audiobooks, brand voice consistency, and accessibility applications

What you can produce with MiniMax Audio

  • Natural-sounding speech audio in 50+ languages generated from a plain-text input
  • Cloned voice profile created from a 10-second reference recording, usable via API
  • Full AI-produced music track up to 5 minutes with vocals, instrumentation, and mixing from tagged lyrics
  • Low-latency sub-250ms streaming speech output for real-time voice agent integration
  • Noise-cleaned vocal-isolated audio output via the Voice Isolator tool
  • Batch long-form synthesis of up to 10 million characters per request for audiobook pipelines
Advertisement

ASEAN Perspective

MiniMax Audio in Southeast Asia

MiniMax Audio supports Vietnamese and Thai among its 50+ languages — two ASEAN languages with complex tonal structures where many Western TTS platforms struggle — and the MiniMax-Speech technical report specifically highlights these as benchmarked languages. Indonesian is also in the supported set, covering the region's two largest internet markets. However, dedicated ASEAN-language voice libraries are thinner than those for English, Mandarin, and Japanese, and sub-regional languages like Tagalog and Burmese are not prominently documented. For ASEAN developers building multilingual pipelines, MiniMax Audio offers a cost-competitive API with solid foundational support, but production quality for less-resourced regional languages should be validated with real samples before deployment.

RECATOOLS Verdict

MiniMax Audio is one of the strongest multilingual TTS platforms available as of mid-2026, particularly for East and Southeast Asian languages where Western competitors tend to underperform. The Speech model family topped the TTS Arena leaderboard in 2025, voice cloning works from just 10 seconds of audio, and the API pricing (from roughly $0.04 per 1,000 characters) undercuts ElevenLabs significantly at scale. MiniMax Music 2.5 adds a genuinely capable music generation layer with paragraph-level structural control and natural vocal synthesis that set it apart from earlier generative music tools.

Caveats are real, however. The English voice variety and emotional nuance still lag behind ElevenLabs in side-by-side tests. API documentation is less polished than established Western peers, and the platform's Chinese-first heritage means English support content and community resources are thinner. Music generation lacks stem export and in-session editing, and the free Music tier is effectively trial-only. Businesses requiring guaranteed commercial licensing should confirm terms directly, as licensing language is not prominently displayed. Copyright litigation filed by Disney, Universal, and Warner Bros. Discovery in 2025 is an open risk factor worth monitoring for enterprise buyers.

Independent AI-assisted assessment by RECATOOLS.

What people say

Speech-02 topped the Hugging Face TTS Arena in 2025; by mid-2026 it has slipped to roughly seventh or eighth place as Vocu V3.0, Sonic-series, and Gemini Flash TTS models pushed ahead. That drop matters less than it sounds — reviewers still single out MiniMax for East and Southeast Asian language naturalness, an area where Western TTS vendors routinely stumble.

Six subscription tiers run from $5/month (100K audio points) up to a $999/month Business plan with 20M points and 800 voice slots, plus custom enterprise pricing above that. API rates land around $60–100 per million characters, and voice cloning from a 10-second sample costs about $1.50 per clone — both comfortably undercut ElevenLabs at volume. Music 2.5 and 2.6 extended the platform into full-track generation with section-tag structural control, though there's still no stem export or in-session editing, and the free Music tier is trial-only.

English voice variety and API documentation lag the established Western players, and support content skews toward a Chinese-first audience. The bigger overhang for enterprise buyers: Disney, Universal, and Warner Bros. Discovery are suing MiniMax over its Hailuo video/image product, not the audio tools specifically, but the litigation (MiniMax lost its motion to dismiss in May 2026) is company-level risk worth tracking before signing a contract.

Summary of public user & expert reviews, compiled by RECATOOLS.

Notable facts

  • MiniMax listed on the Hong Kong Stock Exchange on January 9 2026, surging 43% on debut to reach a valuation of $9.3 billion — one of the largest Chinese AI IPOs in recent years.
  • The MiniMax Speech model is an Autoregressive Transformer that topped the public TTS Arena leaderboard in 2025, outperforming ElevenLabs Multilingual v2 on cross-lingual synthesis and tonal language accuracy.
  • MiniMax Music 2.5 understands 14 structural section tags, giving users director-level control over song architecture before a single note is generated.
  • The company name is a direct reference to the minimax decision algorithm from game theory, and all three founders came from SenseTime, one of China's largest computer vision firms.

Frequently asked questions

How short can the reference audio be for voice cloning?
MiniMax Audio can clone a voice from as little as 10 seconds of reference audio (the API documentation specifies a minimum of 10 seconds for the source file, with an optional shorter example clip under 8 seconds for fine-tuning). Quality improves with longer, cleaner recordings.
Does the free tier allow commercial use?
The free tier (approximately 100,000 characters per month) is intended for personal use and testing. Commercial use rights require a paid plan. Check the current terms on minimax.io before any client or revenue-generating deployment, as licensing language is not prominently surfaced in the UI.
What is MiniMax Music 2.5 and how does it differ from the TTS product?
MiniMax Music 2.5 (launched January 28 2026) is a full-track AI music generator — you provide tagged lyrics and a style prompt, and it returns a produced song with vocals, instrumentation, and mixing. It is a separate product from the TTS speech synthesis tools; TTS converts text to spoken spoken voice, while Music 2.5 generates singing with accompaniment.

About this listing

Researched on
Published on
Last reviewed

This entry was compiled from publicly available data including MiniMax Audio's official website, press releases, documentation, and reputable third-party publications. RECATOOLS is not affiliated with MiniMax Audio unless explicitly stated.

Data accuracy

Third-party AI tools update their pricing, features, availability, and policies frequently. Information here may be outdated by the time you read this — we make reasonable efforts to keep listings current, but cannot guarantee absolute accuracy.

For the latest details, please refer to MiniMax Audio directly →

Spotted something out of date? Suggest an update →

Advertisement