Smallest.ai

Small, fast speech models: sub-100ms TTS and cheap multilingual STT

Video & Audio Freemium Has API
Researched · Published · Reviewed
RECATOOLS Score
7.3 / 10
Capability
7
Value for money
8
Ease of use
7
ASEAN readiness
7
API quality
8
Founded
HQ
Users
Launched
Developer

Overview

San Francisco-based Smallest.ai skips the giant-model approach to speech AI, selling sub-100ms Lightning TTS, 38-language Pulse STT and five-second voice cloning through a usage-priced Waves API built for developers wiring voice into real-time agents and telephony.

Advertisement

Pricing

Pricing shown for reference only. These figures reflect RECATOOLS research as of 12 Jul 2026 and may be out of date or incomplete. This is not financial or purchasing advice — always confirm the current price on the provider’s official website before making any decision.

Pay as You Go
Usage-based
API billing per character (TTS) or per minute (STT), no monthly minimum
  • Lightning V3.1 TTS from $0.175/10K characters
  • Pulse STT from $0.003/minute
  • 15 concurrent streams, 100 RPM cap
  • Email + community support
Enterprise
Custom
For production workloads needing SLAs and compliance
  • 99.99% uptime SLA
  • Professional voice cloning + priority support
  • On-premise deployment option
  • HIPAA, SOC 2, SSO/RBAC compliance

What you can produce with Smallest.ai

  • Lightning TTS: sub-100ms streaming latency
  • Pulse STT across 38 languages
  • Instant voice cloning from ~5 seconds of audio
  • Waves REST API with official Python and Node SDKs
  • MCP server for Claude Code / Cursor integration
  • SOC 2, ISO 27001, GDPR and HIPAA compliance (enterprise tier)
  • On-premise / enterprise deployment option
Advertisement

ASEAN Perspective

Smallest.ai in Southeast Asia

ASEAN-region availability and pricing notes coming soon. Drop the editorial team a note via /contact/ if you can supply local context (Singapore/Malaysia/Indonesia/Thailand/Vietnam).

RECATOOLS Verdict

Smallest.ai's pitch is refreshingly narrow: instead of one more general LLM, it ships small, fast models tuned for real-time voice — Lightning TTS clocks sub-100ms latency and Pulse STT covers 38 languages, both billed per minute or per 10K characters rather than a flat subscription. That usage pricing, starting under $0.20 per 10K TTS characters, undercuts a lot of the field and makes it cheap to prototype before committing.

The catch is scale and reputation: an eight-figure seed round and a two-year-old team can't match the voice libraries or brand trust of Google, ElevenLabs or Deepgram, and support outside the paid tiers is community-only. For builders who need a fast, cheap TTS/STT layer inside an agent or telephony stack, it's a legitimate pick — just budget time to validate voice quality against bigger names before shipping to production.

Independent AI-assisted assessment by RECATOOLS.

What people say

Smallest.ai barely registers on G2 or Capterra yet, so most of what's publicly known about how it performs in practice comes from its own benchmark posts and its Product Hunt launch rather than a deep bench of verified customer reviews — worth flagging plainly rather than dressing up as broader validation.

What is there points the same direction consistently. On Product Hunt, launch comments singled out speed specifically: one reviewer wrote that the latency was "so good" and rated it "so much better than Deepgram in terms of performance and latency." The company's own head-to-head posts against Deepgram and ElevenLabs make the same claim with numbers attached — in cross-region latency testing, Smallest.ai reportedly returned the highest share of lowest-latency responses of the platforms compared, which lines up with its positioning: optimize for speed and cost at scale rather than for the most expressive or emotive voice, which is where ElevenLabs still leads.

On the product side, voice cloning needs as little as five seconds of reference audio and covers 50+ languages and accents, and the underlying Lightning model is small enough to run under 1GB of VRAM — relevant if you're evaluating on-device or edge deployment rather than pure API calls. Pricing is transparent and usage-based (from roughly $0.175 per 10,000 characters of TTS and $0.003 per minute of STT), which is unusual in a category where a lot of vendors push you toward a sales call before you see a number.

The honest gap: no independent third-party review site has enough volume yet to tell you how the product holds up at sustained production scale, under real network conditions, across accents outside its strongest languages, or when something breaks and you need more than community support. Founded in 2023 by Sudarshan Kamath and Akshat Mandloi with roughly $8.3M raised from LetsVenture, Sierra Ventures and 3one4 Capital, this is still a young, resource-constrained team competing against far better-funded incumbents — the latency claims look real, but they come mostly from the vendor's own comparisons and a handful of early adopters rather than a mature review base.

Summary of public user & expert reviews, compiled by RECATOOLS.

About this listing

Researched on
Published on
Last reviewed

This entry was compiled from publicly available data including Smallest.ai's official website, press releases, documentation, and reputable third-party publications. RECATOOLS is not affiliated with Smallest.ai unless explicitly stated.

Data accuracy

Third-party AI tools update their pricing, features, availability, and policies frequently. Information here may be outdated by the time you read this — we make reasonable efforts to keep listings current, but cannot guarantee absolute accuracy.

For the latest details, please refer to Smallest.ai directly →

Spotted something out of date? Suggest an update →

Advertisement