Inception Labs

The lab behind Mercury, the first commercially available diffusion LLM, generating text and code at up to ~1000 tokens per second.

LLMs & Chat Paid Has API
Researched · Published · Reviewed
RECATOOLS Score
7 / 10
Founded
2024
HQ
Palo Alto, USA
Launched
Feb 2025
Developer
Stefano Ermon, Aditya Grover, Volodymyr Kuleshov

Overview

Inception Labs is a Palo Alto AI lab founded in 2024 by Stefano Ermon, Aditya Grover, and Volodymyr Kuleshov. It built Mercury, launched February 2025 as the first commercially available diffusion large language model, which generates in parallel rather than token-by-token for major speed gains. Mercury 2 adds a 128K context window, native tool use, and tunable reasoning. Inception raised a $50 million round led by Menlo Ventures in 2025.

Advertisement

Pricing

Pricing shown for reference only. These figures reflect RECATOOLS research as of 3 Sep 2026 and may be out of date or incomplete. This is not financial or purchasing advice — always confirm the current price on the provider’s official website before making any decision.

What you can produce with Inception Labs

  • Text and code generated at over 1,000 tokens per second — roughly ten times typical autoregressive throughput — from the first commercially available diffusion LLM
  • Latency-sensitive agent and code workloads that autoregressive models make feel slow
  • A 128K context window, tool use and tunable reasoning, all added in Mercury 2 over the more limited v1
  • Structured output and translation, where it holds its own rather than trailing
  • ⚠️ A speed-for-quality trade, made explicitly: independent benchmarks place Mercury 2 in the Claude Haiku / Gemini Flash class rather than the Opus or GPT-4 tier, with harder reasoning trailing by roughly 5–15%
  • ⚠️ Vendor benchmark claims worth discounting until third parties replicate them, a v1 with known hallucination and context problems behind it, and a small lab with a lot still to prove
Advertisement
RECATOOLS Verdict

What this is for: Diffusion-based LLMs (the Mercury family) that generate text and code in parallel, reaching roughly 1000 tokens per second, far faster than typical autoregressive models.

Who this is for: Developers needing very low-latency text and code generation who are willing to bet on the diffusion-LLM approach.

Availability: Paid, usage-based API; US-based lab. Distinct from the G42/Jais 'Inception' name; this is the diffusion-LLM company founded by Stefano Ermon.

Independent AI-assisted assessment by RECATOOLS.

What people say

Inception Labs gets attention as the team that actually shipped a commercial diffusion LLM rather than a research demo, and the recurring praise is speed: Mercury 2 generates at over 1,000 tokens per second against roughly 100 for typical autoregressive models, Coverage in The New Stack and developer writeups calls this speed useful for latency-sensitive agent and code workloads. The parallel-generation approach is a good fit for GPU architecture. Over the more limited v1, Mercury 2 adds a 128K context window, tool use, and tunable reasoning.

The model trades quality for speed, a point reviewers consistently make. Independent benchmarks put Mercury 2 in the Claude Haiku / Gemini Flash class, not the Opus or GPT-4 tier. Its quality on harder reasoning tasks trails by roughly 5-15%, though it holds its own on structured output and translation. Take Inception's own benchmark claims with a grain of salt pending third-party replication. It's worth remembering that Mercury v1 had known issues with hallucination and context limits. Practitioner sentiment boils down to cautious interest. Mercury is fast and cheap for the right jobs, but it isn't a frontier-quality replacement yet, and Inception is still a small lab with a lot to prove.

Summary of public user & expert reviews, compiled by RECATOOLS.

About this listing

Researched on
Published on
Last reviewed

This entry was compiled from publicly available data including Inception Labs's official website, press releases, documentation, and reputable third-party publications. RECATOOLS is not affiliated with Inception Labs unless explicitly stated.

Data accuracy

Third-party AI tools update their pricing, features, availability, and policies frequently. Information here may be outdated by the time you read this — we make reasonable efforts to keep listings current, but cannot guarantee absolute accuracy.

For the latest details, please refer to Inception Labs directly →

Spotted something out of date? Suggest an update →

Advertisement