SAN FRANCISCO, 14 AUG 2026 — The deadline is gone. On 10 August Anthropic made Claude Sonnet 5's introductory US$2 / US$10 rate permanent, and its pricing page now states plainly that the scheduled 1 September rise to US$3 / US$15 "will not occur".

We told readers that deadline was real. It was, when we published. Below: what changed, what we got wrong downstream of it, and why a price scheduled to rise was fixed in place while the cost of producing it climbed.

Our own correction first

On 25 July we published a comparison of eleven frontier models which stated, correctly at the time, that Sonnet 5's rate expired on 31 August. That article now carries three dated corrections.

The worse problem was downstream. Our rate card feeds the public price calculators, and it carried the footnote "Introductory pricing to 31 Aug 2026; US$3 / US$15 from 1 Sep" into the tools. Anyone modelling September costs on this site between 10 and 14 August was told a 50 per cent increase was coming that had already been cancelled. That is fixed, the card is re-verified against the provider's page as of today, and the calculators no longer quote a dead deadline.

One further disclosure, since it is the same relationship in both directions: RECATOOLS articles are drafted with Claude. Anthropic is our supplier, and this is a story about our supplier's prices.

Where the Anthropic card stands today

Sonnet 5 — US$2 / US$10Now the standard rate. Cache reads US$0.20. Batch US$1 / US$5.
Opus 5 — US$5 / US$25Unchanged since launch, and still exactly half Fable 5's US$10 / US$50.
Haiku 4.5 — US$1 / US$5The floor of the range, on a 200k context rather than 1M.
Sonnet 4.6 — US$3 / US$15The rate Sonnet 5 was going to rise to. Its predecessor still charges it.

That last row makes the decision legible. Sonnet 5 was scheduled to converge on what Sonnet 4.6 costs. Instead the newer, more capable model stays a third cheaper than the model it replaced, permanently.

The cost base moved the other way

A price cut made permanent is unremarkable in a falling-cost industry. This is not one, on the evidence of the past fortnight.

What we reportedDirection
OVHcloud raising prices up to 87 per centRAM up sixfold in a year, forecast to twelve times, with normalisation hoped for in 2029.
Foxconn's Q2Its own chief executive names TSMC's CoWoS packaging, not factory space, as the 2027 ceiling on AI server volume.
Anthropic's Riot leaseUS$9.1bn for 191MW over 20 years, plus a reported US$10bn with Volta Infra for Norwegian capacity.

Inference runs on memory, packaging capacity and electricity. Every one of them got dearer or scarcer in the month a per-token price was fixed downward.

Somebody is absorbing that gap. The three likeliest explanations are not mutually exclusive.

The first is efficiency: serving cost per token falls faster than input costs rise, through better utilisation, quantisation, batching and caching. Plausible, unmeasurable from outside, and the reason cache reads at one tenth of base input exist at all.

The second is competition. GPT-5.6 Luna sits at US$1 / US$6 and Gemini 3.1 Pro at US$2 / US$12 below 200k tokens. A US$3 input rate would have priced Sonnet 5 above both. Withdrawing the rise reads as arithmetic rather than generosity.

The third is that this is a land-grab funded by capital rather than margin, at a company reported to be preparing for a public listing. Ali Ghodsi's line about the current moment — that the world remains largely unchanged "except that token spending is rising" — describes the strategy from the buyer's side. Where volume is the thesis, price becomes the instrument for buying it.

The 30 per cent nobody prices in

A detail in Anthropic's pricing documentation changes every cross-generation comparison, ours included.

Claude 4.7 and later models use a newer tokenizer. It produces approximately 30 per cent more tokens for the same text, and the documentation is explicit that the exact increase depends on content and workload shape.

The effect on a bill is straightforward: a per-token rate is a price per unit, and Anthropic changed the size of the unit. If the same document now becomes 30 per cent more tokens, then a headline rate of US$2 per million behaves, for that document, like roughly US$2.60 against a model on the older tokenizer. The cut is real, but smaller than the headline rate suggests.

This does not apply within a generation — Opus 5 against Fable 5 is a like-for-like comparison, and the half-price finding from July stands. It applies whenever a table puts a 4.7-or-later model beside anything older, which is what nearly every price comparison on the internet does, ours included.

This might be the single most under-reported number in frontier model pricing. Anthropic discloses it in a footnote, but we have not seen it factored into any published comparison.

The price of a token is no longer the price of the work

Four line items have appeared alongside the per-token rates since our July piece, and none of them is a token.

Fast mode charges US$10 / US$50 for Opus 5 and Opus 4.8, double the standard rate, for faster output in research preview. Data residency multiplies every pricing category by 1.1 when inference is pinned to the United States. Beyond those two, Managed Agents adds US$0.08 per session-hour on top of whatever the tokens cost, while code execution comes free alongside web search or fetch and is otherwise billed by container-hour past 1,550 free hours a month.

Individually these are minor. Together they mean a cost model built on token rates alone is incomplete — and always in the provider's favour. An agent workload with US-pinned inference, fast mode and a long-running session pays a materially different rate than the table implies.

Why the thread exists

The gap between this entry and the last one is the reason for keeping a running thread.

Three weeks separate a table that was accurate on publication from a table with three corrections in it. Nothing about the July piece was careless — the deadline was announced by the provider, printed on its own page, and true when we read it. It stopped being true on a Monday in August, quietly, in a footnote.

Per-token pricing is published, precise and easy to tabulate. It also moves faster than anyone re-reads it. A price comparison is a claim with a decay rate rather than a fact you establish once. This one decayed in under a month.

This is the argument for structured data with dates on it, rather than prose. Every number in this article comes from a rate card that records where it was read and when, so a wrong figure is a sourcing failure with a timestamp rather than an unattributable error — and re-verifying it is a task somebody can actually do. The discipline surfaced this correction. It did not prevent it, and it was never going to.

What to do with this

If you built a September budget on US$3 / US$15, rebuild it; the change is in your favour.

Then check which tokenizer your models are on. If your comparison spans Claude 4.6 and 4.7, you are comparing units of different sizes, and the older model is cheaper than your spreadsheet says relative to the newer one.

And itemise the non-token charges. Session-hours, residency multipliers and speed premiums do not appear in any price comparison table, ours included, and they are where the difference between a modelled bill and an actual one now lives.

What to watch

Whether any other provider follows Sonnet 5 down. One lab making a permanent cut is a pricing decision. A second one doing it would be a market.

Whether the physical costs reach the token price. Memory at twelve times and packaging capped through 2027 have to surface somewhere, and so far every visible adjustment has been absorbed rather than passed on.

And whether the tokenizer difference ever gets stated in a comparison other than this one — a disclosed fact that quietly undoes a third of the arithmetic everybody is doing.