StartGemini 3.6 Flash Is Out. Google Still Doesn't Have a Model in the Top Ten.
By Addy · July 21, 2026
Google released three Gemini models on July 21: Gemini 3.6 Flash, Gemini 3.5 Flash-Lite, and Gemini 3.5 Flash Cyber. All three solve recognizable production problems. None of them solves the problem hanging over Google's model lineup.
The flagship is still missing.
Artificial Analysis measured Gemini 3.6 Flash at 304 output tokens per second, more than twice as fast per task as its predecessor and about 18% cheaper per task. It also gave the model exactly the same Intelligence Index score as Gemini 3.5 Flash: 50. The model page places it 21st among 187 comparable models.
That makes 3.6 Flash an engineering improvement, not an intelligence breakthrough. Google made the same measured capability faster, less verbose, and cheaper to run. It did not move into the group now setting the frontier.
The launch improves the economics of Gemini. It does not move Google's intelligence ceiling.
Flash Got Better at Being Flash
The strongest case for Gemini 3.6 Flash is operational. It uses 17% fewer output tokens than 3.5 Flash on the Artificial Analysis evaluation suite. Its average time per task fell from 2.7 minutes to 1.3 minutes, helped by faster generation and fewer reasoning steps. Its measured cost per task dropped from USD 0.59 to USD 0.50.
Google also cut output pricing from USD 9 to USD 7.50 per million tokens while leaving input at USD 1.50. The model retains a 1 million token context window, multimodal input, and configurable reasoning. For a company running millions of extraction, coding, document-analysis, or tool-use requests, those changes compound into real money and real latency.
Google's own results show narrower capability improvements too. DeepSWE rose from 37% to 49%, OSWorld-Verified from 78.4% to 83%, and GDPval-AA v2 from 1349 to 1421. Artificial Analysis measured that last improvement independently.
The composite score still did not move. Gemini 3.6 Flash scored 50, matching 3.5 Flash. Its Humanity's Last Exam result slipped three points even as its agentic and knowledge-work scores improved. That is what a flat composite can hide: the model became better for some production jobs without becoming more capable in the round.
This is not a trivial release. Halving task time at nearly the same measured intelligence is valuable. It is also categorically different from releasing a model that changes who leads.
The Top of the Table Moved Without Google
The frontier expanded rapidly before Gemini 3.6 Flash arrived. Claude Fable 5 leads the Artificial Analysis Intelligence Index at 60. GPT-5.6 Sol follows at 59, Kimi K3 at 57, Claude Opus 4.8 at 56, GPT-5.6 Terra at 55, and Grok 4.5 at 54. GPT-5.6 Luna, GLM 5.2, and Muse Spark 1.1 each score 51.
Gemini 3.6 Flash scores 50.
Artificial Analysis now counts six labs with at least one model above 50: Anthropic, OpenAI, Moonshot AI, SpaceXAI, Z AI, and Meta. Google is not one of them. A one-point gap is small in capability terms and large in what it says about the current lineup. Every lab in that group has shipped something the independent composite rates above Google's newest public model.
The comparison requires one important caveat. Flash is designed as Google's fast, economical workhorse, not its maximum-capability flagship. Judging it against Fable 5 or GPT-5.6 Sol is like judging a production sedan against race cars because they appear on the same speed chart. Google is optimizing a different point on the cost-latency curve.
That defense is valid. It also leads directly to the question the launch cannot answer: where is the model Google built for the race-car category?
Flash-Lite Is the Cleaner Upgrade
Gemini 3.5 Flash-Lite makes a more straightforward generational jump. Artificial Analysis scores it at 36, up 11 points from Gemini 3.1 Flash-Lite. Average task time falls from one minute to 0.6 minutes, and output speed reaches 350 tokens per second.
It is priced at USD 0.30 per million input tokens and USD 2.50 per million output tokens. That makes it attractive for high-volume agentic workflows, translation, search, document processing, and workloads where the cost of a frontier model would be wasteful.
The cheap-model label needs context. Flash-Lite's measured cost per task rose from USD 0.04 to USD 0.09 because Google increased token pricing relative to 3.1 Flash-Lite. It is faster, substantially smarter, and still inexpensive in absolute terms, but it is not cheaper than the model it replaces on the Artificial Analysis task mix.
That trade is easier to defend than a flat intelligence score. Google is charging more for a model that independently measures 11 points better and completes work faster.
Flash Cyber Is the Most Aggressive Release and the Least Testable
Gemini 3.5 Flash Cyber is a specialist model built on 3.5 Flash and fine-tuned to find, validate, and patch software vulnerabilities through Google's CodeMender system. The premise is compelling: a lightweight model can search more code paths for the same budget than a large general-purpose model, and several agents can combine their findings into one report.
Google says Flash Cyber found 55 unique confirmed issues in a fixed set of V8 JavaScript engine runs, compared with 47 for mainline 3.5 Flash and 36 for Claude Opus 4.6. It is already being used inside Google across Chrome, Android, Cloud, Ads, and YouTube.
Those are Google-run evaluations, not independent results. Google's CyberGym comparison also relies partly on provider-reported competitor scores, and its production scanning evaluation omits newer competing models that refused the tasks because of their safety controls. The results are interesting, but the harness and access policy matter as much as the bars.
Most developers cannot resolve that uncertainty themselves. Flash Cyber will be available only to governments and trusted partners through a limited CodeMender pilot, with broader expansion promised later. Google's most strategically unusual model in this launch is also the one the public cannot independently test.
The Model Google Needs Is Still in Testing
Gemini 3.5 Pro was supposed to be Google's maximum-capability release. Sundar Pichai said at Google I/O that it was due in June. June passed without it.
Bloomberg reported on July 16 that the model was months behind schedule while Google worked to improve capabilities, particularly coding. Google's launch post now says 3.5 Pro is testing with partners and will become broadly available as soon as it is ready. It provides no replacement date.
That wording matters. Google did not quietly cancel the model, and partner testing means it exists beyond a research checkpoint. But as of July 21, developers cannot use it, independent evaluators cannot measure it, and the public leaderboard cannot give Google credit for what it might become.
Meanwhile, the frontier moved. Fable 5 took the lead. GPT-5.6 arrived one point behind it. Kimi K3 entered third. Grok 4.5, Muse Spark 1.1, and GLM 5.2 turned what had been a two-lab contest into a six-lab cluster in roughly six weeks.
Google's response is not only 3.5 Pro. The company says it has begun Gemini 4, its most ambitious pre-training run yet. That may eventually reset the argument. Today it is a training run, not a model endpoint.
Efficiency Is a Strategy, Not a Leaderboard Excuse
Google's emphasis on throughput is rational. Most production traffic does not require the smartest available model. It requires a model that is good enough, predictable, fast, multimodal, and cheap enough to call repeatedly. Google has enormous distribution through Search, Workspace, Cloud, Android, and the Gemini app. Improving cost per useful task can matter more to that business than winning one composite benchmark.
Gemini 3.6 Flash is evidence for that strategy. It finishes the same evaluation workload in less than half the time, consumes fewer tokens, and lowers measured task cost without losing its composite score. If the job is processing documents or running a tool loop at scale, those gains may be more valuable than five extra Index points.
But a strategy can be sensible and still leave a gap. Google markets Gemini as a frontier family. Its newest broadly available model sits outside the current top ten, its intended flagship missed its announced month, and its next-generation answer is still in pre-training.
Three models shipped. One became a much better workhorse. One became a stronger high-volume model. One opened a credible cybersecurity direction behind restricted access.
The model that would show Google can still lead the general intelligence race did not ship.
That is the honest reading of Gemini 3.6 Flash: twice as fast, cheaper per task, useful at scale, and exactly as intelligent as the model it replaced on the independent composite.
Google improved the product. The leaderboard is still waiting for the flagship.
Sources:
- Introducing Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber -- Google
- Gemini 3.6 Flash and Gemini 3.5 Flash-Lite: Halving Time per Task -- Artificial Analysis
- Google Gemini Launch Delayed as Tech Falls Short of Goals -- Bloomberg News
Previously on TheQuery: