Fish Audio
TTS and voice cloning with #1 leaderboard ranks and 10-second clones
Overview
Fish Audio is a text-to-speech and voice-cloning platform built on the open-source Fish-Speech/OpenAudio research, offering emotion-controllable speech in 80+ languages, clones from as little as 10 seconds of audio, and a developer API priced well below rivals like ElevenLabs.
Pricing
Pricing shown for reference only. These figures reflect RECATOOLS research as of 3 Sep 2026 and may be out of date or incomplete. This is not financial or purchasing advice — always confirm the current price on the provider’s official website before making any decision.
- Up to 7 min generation
- 3 public voice slots
- Enhanced voice cloning
- Up to 200 min generation
- 10 private voice slots
- Voice Design access
- Up to 1,620 min generation
- Unlimited voice slots
- 5 professional voice slots
- Up to 6,250 min generation
- 10 team seats
- 15 professional voice slots
- Organization-level controls
- SOC 2 compliance
- Custom SSO (coming soon)
What you can produce with Fish Audio
- Voice cloning from 10-15 seconds of reference audio
- Emotion-controllable TTS across 80+ languages
- Cross-lingual cloning (clone in one language, speak another)
- Developer REST API (~$15 per million UTF-8 bytes, S2-Pro)
- Open-source Fish-Speech/OpenAudio models on GitHub
- #1 TTS-Arena2 ranking (S1) and top open-weight model on Artificial Analysis (S2 Pro)
- Commercial-use licensing on paid plans
- Zero-data-retention and on-premise options (Enterprise)
Fish Audio's credibility comes from its open-source lineage rather than marketing copy: the underlying Fish-Speech models on GitHub have over 21,000 stars, and the S1 model topped the TTS-Arena2 leaderboard while S2 Pro ranks as the highest-scoring open-weight model on Artificial Analysis's Speech Arena. That's a genuinely strong technical position in a crowded field.
Practically, cloning from 10-15 seconds of audio is fast to set up, and the API is aggressively priced — roughly $15 per million UTF-8 bytes on the S2-Pro model, which reviewers peg at around 11x cheaper than ElevenLabs for comparable output. The free tier is real but strictly non-commercial; anyone shipping cloned voices commercially needs a paid plan.
Caveats worth knowing: clean source audio matters a lot — background noise in a cloning sample produces audible artifacts — and documentation is lighter than the more polished enterprise TTS vendors. Good fit for developers and creators who want flexible, inexpensive, API-driven voice generation and don't mind an open-source-flavored product over a white-glove enterprise one.
What people say
Fish Audio's strongest credential is technical rather than anecdotal: its S1 model holds the #1 spot on TTS-Arena2 — the most widely cited blind-listening leaderboard for text-to-speech — and the newer S2 Pro model ranks as the top open-weight system (11th overall) on Artificial Analysis's Speech Arena leaderboard, ahead of most other open models and closing in on proprietary ones. Voice-clone similarity scores around 8.8/10 in third-party testing, near the top of the category.
Reviewer feedback (Product Hunt and independent AI-tool review sites) is generally positive on core quality: natural-sounding output, fast generation, and clones that reviewers describe as strikingly close to the source voice from as little as 10-15 seconds of reference audio. The main friction points are around the edges rather than the core product — the free tier's limited credits and some demo features that appear available but don't fully work, plus documentation that trails more polished enterprise competitors. Clean source audio matters: background noise in a training clip produces noticeable artifacts in the clone.
The open-source angle is a real differentiator, not just a marketing line — Fish Audio (originally Fish-Speech, since folded into the OpenAudio research series) maintains fishaudio/fish-speech on GitHub with more than 21,000 stars and a voice library the company says has grown past two million user-contributed voices. Founder Shijia Liao positions the company around that open-research lineage rather than a closed commercial model.
On price, the developer API runs about $15 per million UTF-8 bytes on the S2-Pro model — reviewers commonly cite this as roughly 11x cheaper than ElevenLabs for similar output, a meaningful factor for anyone building TTS into a product at volume. Consumer plans range from a non-commercial free tier (8,000 credits/month) through Plus at $11/month, Pro at $75/month (the most popular tier, three team seats), up to Max at $749/month for heavy usage, with custom Enterprise pricing for zero-data-retention and on-premise needs.
Net: strong on measured quality and price, thinner on long-form enterprise-buyer reviews (G2/Capterra) compared with incumbents like ElevenLabs or Murf.
Summary of public user & expert reviews, compiled by RECATOOLS.
About this listing
This entry was compiled from publicly available data including Fish Audio's official website, press releases, documentation, and reputable third-party publications. RECATOOLS is not affiliated with Fish Audio unless explicitly stated.
Third-party AI tools update their pricing, features, availability, and policies frequently. Information here may be outdated by the time you read this — we make reasonable efforts to keep listings current, but cannot guarantee absolute accuracy.
For the latest details, please refer to Fish Audio directly →
Spotted something out of date? Suggest an update →
More in Video & Audio