HANGZHOU, 14 AUG 2026 — Everyone is cutting except DeepSeek. Within days its V4 models move to peak and off-peak rates, with increases running from about 50 per cent to 1,100 per cent depending on model, token type and time of day.

In the same week Google launched a model at US$0.75 per million input tokens, OpenAI cut GPT-5.6 Luna, and Anthropic made a discount permanent. One of these four is not like the others, and its stated reason is the interesting part.

What changes on Sunday

17 hours cheap, 7 dearOff-peak is half the peak rate. Peak runs 01:00–04:00 and 06:00–10:00 UTC.
V4-Flash outputUS$0.28 flat becomes US$0.66 off-peak, US$1.32 peak.
V4-Pro cache-miss inputUS$0.435 becomes US$0.66 off-peak, US$1.32 peak.
50% to 1,100%The full range across models, token types and hours.

DeepSeek says the reason is capacity, not margin. The goal is to "allocate resources more reasonably" and encourage users to "schedule their tasks based on actual usage". This is congestion pricing, and we are not aware of another major model provider charging by the clock.

The effective date is reported two ways and it matters for a billing change. Several outlets give 16 August, one of them specifying 16:00 UTC; others say 17 August, and at least one describes it as "starting Monday", which is the 17th. Sunday and Monday are one day apart and a day of peak-rate traffic is not nothing. Take the date from DeepSeek's own console rather than from any article, this one included.

Read the 1,100 per cent carefully

The headline number is real and it is the least representative figure in the announcement.

That 1,100 per cent lands on cache-hit input — the cheapest token category there is, the one you pay for re-reading context the provider has already processed. A huge multiple of a tiny number is still a tiny number. Quoting the figure without its context makes the change sound like a blanket repricing.

The categories that carry real spend moved far less. Cache-miss input rises 51 to 214 per cent depending on model and hour; output rises 127 to 371 per cent. Those are serious increases. They are not eleven-fold.

What the cache-hit rise does hit is the workload shape that made DeepSeek attractive in the first place. Long-context agents re-reading a large prompt on every turn live almost entirely in cache hits, and they are precisely the users who will feel a multiple rather than a percentage. As Greyhound Research put it, the cache is where the advantage genuinely erodes.

The divergence, in one table

ProviderDirection this month
GoogleGemini 3.7 Flash launched at US$0.75 / US$3.75 — a 50 per cent introductory discount running to year-end.
OpenAICut GPT-5.6 Luna; shipped an Ultrafast mode reaching 750 output tokens a second on Cerebras hardware.
AnthropicWithdrew a scheduled 50 per cent rise, making Sonnet 5's US$2 / US$10 permanent.
DeepSeekRaising, and introducing time-of-day pricing.

The obvious reading is a lost price war. The reality is that all four face the same shortage, but only DeepSeek is passing the cost on to its customers.

Same constraint, opposite response

We have spent the past fortnight documenting the physical inputs to inference getting scarcer: memory at six times and heading for twelve, packaging capacity capping 2027 output, and a US$9.1bn twenty-year lease for 191MW of Texas electricity.

DeepSeek is the one provider whose pricing now reflects that directly. Peak and off-peak rates are for when capacity, not unit cost, is the constraint. It is the model used by airlines and electricity markets — a candid admission that demand is outstripping supply.

The others face the same inputs and are absorbing them, funded by capital rather than margin. Anthropic has an IPO reportedly in preparation and multi-billion-dollar power leases signed. Google and OpenAI are pricing to hold ground in a market where being the default matters more than the unit economics of any given quarter.

They are different answers to who pays for scarcity now: the customer, or the balance sheet.

The precedent in April

This is not just a China story; DeepSeek is not the first provider to raise prices on capacity grounds this year.

Anthropic raised prices in April for substantially the same reason, according to the reporting on this week's change — we have not independently confirmed that comparison. The difference is what happened next: Anthropic's increase came at a moment when it was raising capital and signing long-term supply, and by August it was withdrawing a scheduled rise rather than adding one. The constraint did not go away — it got financed.

The real distinction is not geography, but the balance sheet. A provider that can raise money against future volume can price below cost for as long as investors believe the volume arrives. A provider that cannot has to charge what serving actually costs, and gets described as expensive for doing so.

Which of those is the sustainable position is not obvious. Congestion pricing at least tells customers the truth about scarcity, and a rate that reflects the cost of service is easier to plan against than one that depends on a funding round.

Congestion pricing has consequences nobody has modelled

The clock is the new thing here, and it deserves more attention than the percentages.

Peak hours of 01:00–04:00 and 06:00–10:00 UTC are business hours in Asia. For a team in Singapore, Jakarta or Ho Chi Minh City, 09:00 to 18:00 local sits substantially inside the expensive window; for a team in California it is the middle of the night. The rate card is neutral and its incidence is not.

Batch and asynchronous workloads can be moved, and DeepSeek is explicitly inviting that. Interactive ones cannot. A customer support assistant answers when customers ask, and a coding agent runs when engineers work — neither can be deferred to a cheaper hour without deferring the business it serves.

The effective increase depends on whether your workload has a flexible schedule or a hard deadline — a detail no price table shows. That is a distinction cost models do not currently make, and anyone in this region reading a headline per-token rate should now check it against their own working day.

What this does to the cheap tier

DeepSeek's role in the market has been to set the floor. Its off-peak rates keep some of that — US$0.22 per million cache-miss input on V4-Flash is still inexpensive by any standard — but the floor is now conditional, and conditional floors are harder to plan against.

It also arrives as Google puts Gemini 3.7 Flash at US$0.75 input on an introductory discount to year-end. Those two facts together compress the gap that made a Chinese provider worth the integration effort for a Western or ASEAN buyer, at least during working hours.

The honest caveat is that introductory discounts expire and congestion pricing can be revised. Every figure here is a claim with a date on it. We just published a correction on a withdrawn deadline, and this table is just as provisional.

What to do about it

Work out what fraction of your DeepSeek spend falls inside 01:00–04:00 and 06:00–10:00 UTC. That single number tells you whether this is a 50 per cent event or a 200 per cent one for you, and it is calculable from existing logs rather than estimated.

Then separate the deferrable from the interactive. Anything batch — evaluation runs, bulk extraction, offline classification, index building — should move off peak, and the saving is exactly half. Anything user-facing should be repriced at peak, because that is what it will actually cost.

And check your cache-hit ratio before assuming the headline. An agent that re-reads a long prompt every turn is the profile this change targets hardest, and it is also the profile most teams have drifted into without measuring.

What to watch

Whether anyone follows DeepSeek onto the clock. Time-of-day pricing is either a one-off response to one provider's capacity problem or the beginning of how inference gets sold, and the next provider to try it decides which.

Whether the increases hold past the shortage. Congestion pricing introduced during scarcity has a way of surviving into abundance.

And whether Western providers can keep absorbing. Three of the four are cutting prices into rising input costs, which is a position funded by expectations rather than by margin. That works until the expectations are repriced.