Google ships three new Gemini Flash models built for agent economics, still no 3.5 Pro
Gemini 3.6 Flash cuts output tokens by 17% at $1.50/$7.50 per million, a 350-tokens-per-second 3.5 Flash-Lite targets high-volume subagents, and a government-only 3.5 Flash Cyber runs inside CodeMender — while the long-delayed 3.5 Pro remains in partner testing and Gemini 4 pre-training begins.
Google DeepMind shipped three new Gemini Flash models on Tuesday and didn’t ship the one that matters for narrative reasons: Gemini 3.5 Pro. The release is a study in what a frontier lab announces when its headline model isn’t ready.
Gemini 3.6 Flash is the volume play, priced at $1.50 per million input tokens and $7.50 per million output tokens. According to Tulsee Doshi, senior director of product management on the Gemini team, it produces 17% fewer output tokens than the outgoing 3.5 Flash on the Artificial Analysis Index, which is the pricing metric that actually matters in an agent economy where models call themselves in loops. Benchmarks moved with it: DeepSWE from 37% to 49%, MLE Bench from 49.7% to 63.9%, OSWorld-Verified from 78.4% to 83.0%, GDPval-AA v2 from 1349 to 1421. Harvey and Hebbia are cited as early customers, both firms whose economics depend on cheap inner-loop reasoning.
Gemini 3.5 Flash-Lite is the throughput play at $0.30 in and $2.50 out per million tokens, clocking 350 output tokens per second on the Artificial Analysis Index. Its GDPval-AA v2 score jumped to 1140 from a predecessor’s 642. It’s already rolling into Google Search, where subagent latency is measured in user attention.
The third model, Gemini 3.5 Flash Cyber, is the interesting one. It’s government-only, runs exclusively inside Google’s CodeMender bug-hunting agent, and can be called up to five times per report. On the CyberGym benchmark against the V8 JavaScript engine, SiliconAngle reports it surfaced 55 unique confirmed issues versus 47 for stock 3.5 Flash and 36 for Anthropic’s Claude Opus 4.6, with 10 findings no other model caught. Salesforce, Robinhood, and Palo Alto Networks are early enterprise testers. Google is keeping the model on a short leash, which is itself a statement about where offensive-capable weights sit in 2026.
What’s missing frames what’s present. Gemini 3.5 Pro was last teased for June; Bloomberg reported internal delays tied to performance targets, and CNBC now says it’s in partner testing. Google simultaneously confirmed pre-training has begun on Gemini 4. The Flash cascade is what you release when the flagship slips and you need to keep the developer surface fed.
Sources
- blog.google, Gemini 3.6 Flash, 3.5 Flash-Lite, 3.5 Flash Cyber
- TechCrunch, Google releases three new Gemini models but no 3.5 Pro
- CNBC, Google Gemini Flash targets Mythos rival
- SiliconAngle, Google expands Gemini, keeps bug hunter on a leash
- AI News, Gemini 3.6 Flash targets enterprise agent token costs
— END —