4 SEP 2026 — The line going round is that four frontier AI models launched in 72 hours at the start of September. Count what shipped and the set contains two models described as new, a looser-safeguard variant of one of them, two point releases, a pricing tier and a feature that is not a model at all. The second new model is OpenAI's GPT-6 Astra, which shipped on 3 September into a staged rollout rather than general availability.

Disclosure: RECATOOLS is written with the assistance of Claude, made by Anthropic, whose releases are among those described here.

What actually shipped

On 1 September Anthropic released Claude Fable 5.1 and Claude Mythos 5.1, characterised as a point release offering the same model at two safeguard tiers, with Fable public and Mythos gated to trusted-access programmes. Perplexity shipped Hybrid Compute the same day, a feature that splits processing between local and cloud rather than a model.

On 2 September Google released Gemini 3.8 Flash, described as a new frontier model and the third Flash release in six weeks, alongside Gemini 3.8 Flash Cyber, a variant with looser mitigations gated through its Fairwind Program. Meta released Muse Spark 1.3 and a pricing-tier variant called Contributor.

Astra is the one everybody counted before it shipped, and it has now shipped. OpenAI released it on 3 September to a limited set of organisations, with ChatGPT Plus, Pro, Business and Enterprise users, the API and AWS following over the days after. It meets OpenAI's critical cybersecurity threshold and ships with those capabilities restricted, which we covered when it was announced, and we have since looked at what its benchmarks actually show.

2Releases described as new frontier models
3rdFlash release from Google in six weeks
2Point releases, from Anthropic and Meta
0Shared benchmarks all of them were run on

A version number is not a capability jump

New model against point release is a practical distinction rather than pedantry, because the two mean different things to anyone deciding whether to re-test their systems.

A new model can change behaviour enough to break a prompt that worked, invalidate an evaluation, or shift where a guardrail sits. A point release is usually a smaller delta, which is why labs number them that way and why nobody re-runs an entire test suite for one.

Counting them together makes the release cadence look explosive even though the capability curve does not match it. Three Flash releases from one lab in six weeks is a fast shipping cadence and evidence about deployment velocity, not about how quickly the models are getting better.

The naming makes this harder than it needs to be. Labs are not obliged to reserve whole numbers for capability jumps, and the increments have drifted. A point release from one lab can be a larger change than a whole number from another, and there is no convention a customer can rely on.

The safeguard tier is the real product distinction

The new thing in this cluster is a packaging decision rather than a model. Two labs shipped the same weights twice over, once with normal mitigations and once with looser ones behind an access programme.

Anthropic's Fable and Mythos split is described exactly that way. Google's Flash and Flash Cyber pairing is the same structure. In both cases the gated variant is the one that will do things the public one declines to do.

We wrote yesterday that three labs shipped cyber models and three vetting programmes in one week. This is the packaging consequence of that: the unit being released is now a model plus a safeguard configuration plus an admission decision, and a release-count that treats those as one launch each has stopped describing the thing accurately.

Nobody ran the same test

The benchmark figures published alongside these releases do not compare. Google reported HLE-Verified at 54.9 per cent for Gemini 3.8 Flash and CWE-Bench pass@1 at 47.2 per cent for the Cyber variant, both vendor-run. Anthropic published no benchmark figure with the point release. Meta published none either.

The only common measure across two of them is a third-party index, Artificial Analysis, which puts Gemini 3.8 Flash at 59 and Muse Spark 1.3 at 61 in its highest-effort setting. Two points on a composite index is not a difference anyone should act on, and it does not cover the other releases at all.

So four vendors shipped things in the same week, and no public evaluation covers all of them. Anyone choosing between them is comparing marketing across incompatible tests.

Why the cadence is accelerating anyway

The underlying shift is real even though the count is inflated. Labs have moved from occasional large releases to continuous shipping, with point releases, tier variants and pricing changes arriving between the big ones.

That is what a product organisation does once the research has a delivery pipeline attached, and it has a practical consequence: the model behind an API changes more often than most teams re-test, and a version number that increments by one tenth still arrives in production.

The defensive habit is unglamorous. Pin versions where a provider allows it, keep a small evaluation you can run in an hour, and run it when a number changes rather than when a headline says a frontier moved.

What would make a release count meaningful

Three things, none of which any lab currently publishes together. Whether a release changes weights or only serving configuration. Whether the prior version stays available and for how long. And a benchmark result on a test the lab did not select.

Until those exist, a release tracker is counting announcements. That is a useful thing to count, and it is not the same as counting capability. The difference is why a 72-hour figure sounds like more than it is.

None of which says nothing is happening. Six releases from four companies inside three days is an unusual week by any measure. It is a busy shipping calendar, not four frontiers moving at once.