Google released Gemini 3.6 Flash on 21 July 2026 at US$1.50 per million input tokens and US$7.50 per million output — the same input price as 3.5 Flash, and US$1.50 less on output.

The interesting part is not the cut. It is that Google pairs it with a claim that the model also emits fewer tokens for the same work, and those two things multiply. We priced what that actually means, and where the claim gets shaky.

US$1.50 / US$7.50Gemini 3.6 Flash per 1M tokens, in/out
0%change to the input price — the cut is output-only
16.7%sticker cut on output vs 3.5 Flash
~31%effective cut IF the third-party token claim holds

Only one number moved

Input stays at US$1.50 per million on both models. Context caching is unchanged at US$0.15 per million plus US$1.00 per million tokens per hour of storage. The batch tier moves in step with the standard tier — US$0.75 in and US$3.75 out, against 3.5 Flash's US$4.50.

So this is not a repricing of the tier. It is a targeted cut on the half of the bill that output-heavy work actually feels, and input-heavy workloads — retrieval, long context, tool-calling agents — see nothing from it at all.

The two cuts, and how they multiply

Google's model page states that 3.6 Flash "reduces output token usage by 17% compared to 3.5 Flash, according to Artificial Analysis Index". If that holds on your workload, you are paying a lower rate on fewer units, and the arithmetic compounds:

Computed by RECATOOLS26 July 2026
StepFigureWhere it comes from
3.5 Flash output$9.00 / 1MGoogle’s pricing page
3.6 Flash output$7.50 / 1MGoogle’s pricing page
Sticker cut16.7%computed: (9.00 − 7.50) ÷ 9.00
Claimed token reduction17%Artificial Analysis, quoted by Google
Effective output cost, same work$6.225 / 1Mcomputed: 7.50 × (1 − 0.17)
Effective cut vs 3.5 Flash30.8%computed: (9.00 − 6.225) ÷ 9.00

The second figure is only as good as the token-efficiency claim behind it, which is a third-party measurement — see the caveat below.

A sixth off the sticker becomes roughly a third off the bill for the same work. That is the number worth planning around — and it is also the number most likely to disappoint, because it inherits every assumption in the measurement behind it.

Where the field sits now

Computed by RECATOOLS26 July 2026
ModelIn / Out $/MBlend 3:1Blend 10:1Cached in
Gemini 3.6 Flash$1.50 / $7.50$3.00$2.05$0.15
Gemini 3.5 Flash$1.50 / $9.00$3.38$2.18$0.15
Gemini 3.5 Flash-Lite$0.30 / $2.50$0.85$0.50$0.03
GPT-5.6 Luna$1.00 / $6.00$2.25$1.45$0.10
Claude Haiku 4.5$1.00 / $5.00$2.00$1.36$0.10
Grok 4.3 ¹$1.25 / $2.50$1.56$1.36
DeepSeek V4-Flash ²$0.14 / $0.28$0.18$0.15$0.0028

Blended $/M = (r × input + output) ÷ (r + 1). Prices read from each provider’s official pricing page.

¹ Rate shown for prompts under 200k tokens; doubles at or above.  ² DeepSeek's headline cache-hit rate is far lower again, but cache-miss is the honest basis for comparison.

The cut moves Gemini within the pack, not ahead of it. At an agent-shaped 10:1 blend, 3.6 Flash lands at US$2.05 per million — better than the US$2.18 it replaces, still above Claude Haiku 4.5 and Grok 4.3, and roughly thirteen times DeepSeek V4-Flash. Google's own Flash-Lite undercuts it four to one.

The compounding claim is what would change the ranking. Apply the 17% reduction and 3.6 Flash's effective position improves against every rival that has not also become more concise — which is precisely why the provenance of that 17% matters more than its size.

The caveat that has to travel with the number

The token-efficiency figure is not Google's measurement. It is Artificial Analysis's, quoted by Google, on a benchmark index rather than on your traffic. Verbosity is also the most workload-sensitive property a model has: a model that is 17% more concise on multi-step agent tasks may be no more concise on short classification calls, and could be worse on long-form generation.

Nothing here is dishonest — Google attributes the figure clearly. But a rate card is a contract and a benchmark result is an observation, and only one of them is enforceable. We track Gemini's tiers and pricing history in the directory:

The caveats that matter

  • Input-heavy work sees nothing. Retrieval and long-context pipelines pay the same US$1.50 they did before.
  • Grounding is billed separately. Search and Maps grounding is free for 5,000 prompts a month, then US$14 per 1,000 queries — and is unavailable on the free tier.
  • Google does not date the launch in its docs. Neither the pricing page nor the model page carries a release date; 21 July comes from Google's own blog post and independent coverage the same day.

Key takeaways

  • The cut. Gemini 3.6 Flash costs US$1.50/US$7.50 per million tokens — input unchanged, output down a sixth from 3.5 Flash's US$9.00.
  • The compounding. Combined with a claimed 17% reduction in output tokens, the effective cut on equivalent work is closer to a third.
  • The provenance. That 17% is Artificial Analysis's measurement quoted by Google, on an index — not a guarantee about your workload.
  • The field. At a 10:1 agent blend the model sits at US$2.05 per million: mid-pack, undercut by Claude Haiku 4.5, Grok 4.3, Google's own Flash-Lite and DeepSeek V4-Flash.
  • The asymmetry. Output-heavy workloads capture the whole benefit; input-heavy ones capture none of it.