Google released three new Gemini models on July 21, 2026: 3.6 Flash, 3.5 Flash-Lite and the specialized 3.5 Flash Cyber. All three target cheaper, faster deployment of AI agents at scale. Yet the loudest signal is what Google did not ship — the flagship Gemini 3.5 Pro.
Key takeaways
- Gemini 3.6 Flash uses 17% fewer output tokens than 3.5 Flash (Artificial Analysis Index), with reductions reaching 65% on the DeepSWE benchmark.
- Pricing for 3.6 Flash: $1.50 per 1M input tokens and $7.50 per 1M output tokens.
- Gemini 3.5 Flash-Lite runs at 350 tokens per second and doubles its predecessor on Terminal-Bench 2.1 (54% vs 31%).
- Gemini 3.5 Flash Cyber is fine-tuned for vulnerability detection and available only through a limited pilot for governments and trusted partners.
- Gemini 3.5 Pro was not released — Bloomberg reports it fell short of internal performance targets.
Three models, one goal: cheaper inference
Gemini 3.6 Flash is Google’s new „workhorse?workhorse: A dependable, versatile everyday model — built for volume work, not for topping benchmarks.” — a model for everyday coding, knowledge work and multimodal data.
The core change is economic: fewer tokens consumed at a lower price per output token. On DeepSWE (software engineering tasks) 3.6 Flash reaches 49% versus its predecessor's 37%, and on MLE Bench it climbs to 63.9% from 49.7%. On the OSWorld-Verified agentic benchmark the score rises from 78.4% to 83.0%.
3.5 Flash-Lite competes on a different axis — throughput. At 350 tokens per second and $0.30 per million input tokens, it is the fastest and cheapest model in its class.
The sharpest gains show up in agentic tasks: Terminal-Bench 2.1 jumps from 31% to 54%, and GDPval-AA v2 rises from 642 to 1140 points. The message is that Google is tuning Flash-Lite for multi-tool workflows where cost and latency matter more than raw intelligence.
The third model, 3.5 Flash Cyber, is narrowly specialized. Google fine-tuned it to detect and fix code vulnerabilities, paired it with the CodeMender security agent, and points to frontier-competitive results on the CyberGym benchmark. It is not a general release, however — it reaches only governments and trusted partners through a limited-access pilot.
| Model | Role | Headline result | Price (1M in / out) |
|---|---|---|---|
| Gemini 3.6 Flash | Everyday workhorse | DeepSWE 49% (from 37%) | $1.50 / $7.50 |
| Gemini 3.5 Flash-Lite | Speed and cost | Terminal-Bench 54%, 350 tok/s | $0.30 / $2.50 |
| Gemini 3.5 Flash Cyber | Cybersecurity | Frontier-competitive on CyberGym | Limited pilot |
The missing piece: Gemini 3.5 Pro
The most telling thread concerns the model that isn't here. Flagship Gemini 3.5 Pro was last updated in February, and — as Bloomberg previously reported — Google struggled with delays in hitting internal benchmark targets. Product lead Logan Kilpatrick acknowledged the team is still testing 3.5 Pro with partners.
We're still testing 3.5 Pro with partners and expect it to land soon.
Logan Kilpatrick, Gemini API product lead, Google
At the same time Google confirmed that pre-training?pre-training: The initial, heaviest and most expensive training pass on a massive dataset, before fine-tuning. of Gemini 4 has begun, describing it as its "most ambitious pre-training run yet." That leaves the company shipping a layer of cheaper Flash models while thinking a full generation ahead — and leaving a gap on the most important, flagship shelf.
Competitive context
For comparison, OpenAI and Anthropic have maintained a faster flagship release cadence in recent months. The absence of 3.5 Pro exposes a tension in Google's strategy: it delivers efficient, cheap operational models but delays the one that genuinely competes with Opus 4.8 or GPT-5.6 at the top. The Flash layer solves the cost problem for agentic deployments, but it does not answer where Gemini stands on the capability frontier.
Why it matters
This release shows where model competition actually plays out today — not the highest score on a single benchmark, but the per-unit cost of running agents in production. Token reductions of 17–65% and sub-dollar input pricing on Flash-Lite feed directly into the bills of companies building agentic systems at scale.
For engineering teams, a cheaper token often matters more than a few percentage points on a reasoning benchmark. At the same time, the missing 3.5 Pro is a warning sign: if Google cannot ship its flagship on schedule, it cedes initiative in the segment that shapes who is perceived as the leader. Releasing the Flash layer while announcing Gemini 4 may be an attempt to paper over that gap.
What's next?
- Gemini 3.5 Pro is expected to reach broader availability "soon" — Google gave no firm date, and prior delays (per Bloomberg) make the timeline uncertain.
- Gemini 4 pre-training is underway — the next milestone is the first public data on results or a release date.
- 3.5 Flash-Lite is rolling out in Google Search — its reach and impact on result quality will become visible over the coming weeks.





