MOUNTAIN VIEW, 15 AUG 2026 — Google released Gemini 3.7 Flash on Thursday at US$0.75 per million input tokens and US$3.75 per million output, half what its predecessor cost at launch. The predecessor came out three weeks ago.
Read the second sentence of the pricing page before you budget on the first. The rate is introductory and it expires on 31 December 2026, after which input goes to US$1.50 and output to US$7.50. What is quoted is a time-limited offer rather than a rate.
What it costs, and when that changes
| Now → 31 Dec 2026 | From 1 Jan 2027 | |
|---|---|---|
| Input / M tokens | US$0.75 | US$1.50 |
| Output / M tokens | US$3.75 | US$7.50 |
| Status | Introductory | Doubles |
Four and a half months at the low rate, then a hundred per cent increase on a fixed date. If you are sizing an agent workload that will still be running in January, the number to plan against is the right-hand column.
The benchmarks, and whose they are
The capability gains are large. All figures below are Google's.
DeepSWE moving from 49 to 65 per cent in three weeks is the number that should make you look twice. So is AutomationBench nearly doubling. Google says the result puts 3.7 Flash ahead of Claude Sonnet 5 and GPT-5.6 Terra on coding and agent work.
That claim is a vendor's, measured by the vendor, on benchmarks the vendor selected, which makes it a starting point for enquiry rather than a finding. Google's own results do not show the model displacing higher-priced competitors across the board. The useful signal is narrower: a cheap tier has become competitive at particular kinds of work.
Three weeks is the part that should unsettle you
Gemini 3.6 Flash launched three weeks ago. Anyone who standardised on it, wrote prompts against it, tuned an agent scaffold around its behaviour and signed off a cost model has been undercut by fifty per cent by their own supplier, in under a month.
Teams that have not committed yet get a better deal than the ones that did. Model choice has become a decision with a shelf life measured in weeks, and the engineering cost of switching — re-evaluating prompts, re-running your own regression suite, re-checking output shape — is now frequently larger than the price difference that motivated it.
Do not chase every release. Stop treating the model identifier as a constant in your codebase instead. Where swapping models costs a configuration change and a test run, a three-week cadence works in your favour. Where it costs a refactor, you will pay for it repeatedly.
What the cut says about the last price
Halving the price of a better model three weeks after its predecessor shipped has a quieter implication.
If Gemini 3.6 Flash could be undercut by fifty per cent that quickly, by a model that scores higher on Google's own coding benchmarks, then 3.6 Flash's launch price was never anchored to what it cost to serve. Inference costs do fall, sometimes sharply, but not by half in twenty-one days on hardware that was already deployed.
In practice, published rates in this market carry no information about the cost floor beneath them. They are positions in a competition, and they can move as far and as fast as the competition requires. Anyone building a margin assumption on the idea that a vendor cannot go lower — or will not go higher — is reading the wrong signal.
Two vendors, opposite directions, same week
The expiry date is interesting because a competitor just did the reverse.
Anthropic withdrew its scheduled Sonnet 5 increase last week and made the introductory rate permanent, cancelling a rise that had been set for 1 September. Google has introduced a rate that doubles on 1 January. The two moves landed within days of each other.
Set that beside Anthropic reporting a fourteen-fold revenue rise and DeepSeek raising prices outright, and the shape of the market is legible. Nobody is pricing to cost. Everyone is pricing to a strategic position. Anthropic wants to keep customers ahead of a listing, DeepSeek is rationing scarce capacity, and Google is buying agent developers now to reprice them later.
An introductory rate with a published expiry is the most honest of the three, incidentally. It tells you exactly when the bill changes. The alternative is a permanent-sounding price that moves without notice, which is what most of this market offers.
Why a cheap tier matters more here
The cheap tier is where most regional development actually happens.
Agent workloads consume tokens in a way chat does not. A multi-step agent that plans, calls tools, reads results and retries can burn through a hundred times the tokens of a single question for one unit of work. At that volume the per-token rate stops being a line item and becomes the product's gross margin.
For a team in Singapore, Jakarta or Manila building on a regional price point rather than a Silicon Valley one, the difference between US$0.75 and US$1.50 is the difference between a viable unit economic and a demonstration. The expiry date, then, deserves more attention than the benchmark table. Benchmarks tell you what the model can do. The date tells you how long you can afford to run it.
The advice is short. Build the cost model on US$1.50 and US$7.50. If the product works at those numbers, the introductory rate is four months of margin. If it only works at US$0.75, what you have is a window rather than a business.
What we could not establish
Whether the introductory pricing will actually end on 31 December. Vendors have extended these offers before, and Anthropic just cancelled one, making the date more of a stated intention than a commitment.
Also unestablished are the context window and rate limits at each tier. It is not clear whether the benchmark comparisons used equivalent reasoning settings across models, how "code quality" was measured, or whether the gains hold on non-English workloads, and no independent evaluation has reproduced the coding numbers. No third party had published one at the time of writing.
What to watch
First, whether an independent benchmark confirms the DeepSWE jump. A sixteen-point move in three weeks on a cheap tier is either an important result or a measurement artefact, and only somebody outside Google can tell you which.
Second, whether the 31 December date holds. If it slips, introductory pricing across this market becomes a negotiating posture rather than a schedule, and everyone will plan accordingly.
And finally, whether Gemini 3.8 Flash arrives before January. On the current cadence it would, which would make the expiry date on this model a question about a product nobody is using by then.