StartGemini 3.8 Flash Beats Terra. The Price Is the Point.
By Addy · September 2, 2026 · Editorial standards
Google released Gemini 3.8 Flash on September 2, 2026, three weeks after Gemini 3.7 Flash and only six weeks after 3.6 Flash. The version number moved by one decimal point. The competitive position moved further.
In Google's launch table, Gemini 3.8 Flash beats GPT-5.6 Terra on 13 of the 14 benchmark rows where both have a score. The only loss is Terminal-Bench 4.0. On the independent Artificial Analysis Intelligence Index, Gemini 3.8 Flash at High effort scores 59, ahead of Terra at its Max setting with 57.
The price difference is larger than the score difference. Gemini costs USD 0.75 per million input tokens and USD 3.75 per million output tokens during Google's introductory period. Terra costs USD 2 and USD 12. Google is charging less than half as much while its Flash model now wins most of the comparison built for a much more expensive class.
That is the real release. Gemini 3.8 Flash is no longer interesting because it is cheap for a Flash model. It is interesting because the cheap model has started beating the balanced premium model.
What Google Actually Released
Gemini 3.8 arrives in two versions. Gemini 3.8 Flash is the generally available workhorse for coding, research, professional analysis, computer use, and long-running agent workflows. Gemini 3.8 Flash Cyber uses the same underlying intelligence with more permissive cybersecurity mitigations and is restricted to approved defenders through Google's Fairwind Program.
The public Flash model is available through the Gemini API, Google AI Studio, Android Studio, Google Antigravity, Stitch, Gemini Enterprise, the Gemini app, AI Mode in Search, and Gemini in Google Sheets. Google AI Pro and Ultra subscribers receive it in consumer products.
Google describes 3.8 Flash as its best reasoning and coding model yet, not merely its best Flash model. The company says the gains came partly from training in cybersecurity and from longer agentic AI loops that repeatedly evaluate and refine the model's work.
The one-million-token context window remains part of the Flash proposition. So do adjustable effort levels. The model can spend more reasoning and tool calls on hard work or run at a lower effort when latency and token use matter more.
Gemini Beats Terra on 13 of 14 Rows
Google's benchmark table is unusually direct because it puts Gemini 3.8 Flash beside Gemini 3.7 Flash, Claude Opus 5, Claude Sonnet 5, GPT-5.6 Sol, and GPT-5.6 Terra.
The table does not require a composite score to make the Terra comparison.
Gemini beats Terra on DeepSWE v1.1, 73.7% to 69.6%. It leads Vals Finance Agent v2 by seven points, Harvey's Legal Agent Benchmark by 9.2 points, Terminal-Bench 2.1 by two points, and OSWorld 2.0 by 8.8 points. It also leads on GDPval-AA v2, GDP.PDF, CharXiv Reasoning, LVBench, HLE-Verified, both BioMysteryBench difficulty groups, and LABBench2.
Terra's one win is Terminal-Bench 4.0, where it scores 23.6% against Gemini's 19.1%. Neither is close to GPT-5.6 Sol at 37.3% or Claude Opus 5 at 51.8%. The row matters because Terminal-Bench 4.0 tests general AI agent capability rather than the narrower terminal coding measured by version 2.1. Google wins the coding-oriented terminal test and still has a larger gap on the broader agent test.
The table should still be read as Google's evidence. Effort settings, agent harnesses, tool access, graders, and benchmark versions can move these numbers. A benchmark score measures the configured system that ran the task, not a model name floating above its tools. Google's methodology says Terra and Sonnet use their maximum available reasoning settings where reported results exist. Some rows are self-computed by Google, while others come from public leaderboards or providers' published numbers. The comparison is documented, but it is still assembled by the company launching the model.
It is also not a claim that Gemini 3.8 Flash beats every frontier model. Claude Opus 5 remains ahead on GDPval-AA v2, Terminal-Bench 4.0, OSWorld 2.0, and the human-solvable BioMysteryBench set. GPT-5.6 Sol leads GDP.PDF. On DeepSWE, Opus 5 scores 74.0% against Gemini's 73.7%, effectively the same result at the displayed precision.
The narrower claim survives those caveats: against Terra, in the table Google chose to publish, Gemini wins almost everywhere.
The Jump From Gemini 3.7 Flash Is Real
Gemini 3.7 Flash was already Google's repair job. It moved the independent Intelligence Index from 52 on 3.6 Flash to 56 at High effort, and it closed much of the coding gap with GPT-5.6 Terra. Three weeks later, 3.8 Flash moves the High-effort index again, from 56 to 59.
Artificial Analysis reports 3.8 Flash at 59 on High, 57 on Medium, and 52 on Low. The corresponding 3.7 Flash scores are 56, 53, and 51. The improvement is therefore not evenly distributed. Medium gains four points, High gains three, and Low gains one. Google improved the part of the model that is allowed to work longer more than the setting optimized for minimum effort.
Google's table shows the same shape across individual tasks. DeepSWE rises from 65.3% to 73.7%, an 8.4-point gain. OSWorld rises from 50.6% to 59.0%, also 8.4 points. Terminal-Bench 4.0 moves from 11.2% to 19.1%. On the human-difficult BioMysteryBench set, the score jumps from 43.5% to 56.5%, the largest improvement in the table.
The smaller gains are still consistently positive. HLE-Verified rises 1.3 points, GDP.PDF rises one point, and Harvey's legal benchmark rises 1.2. Gemini 3.8 Flash beats 3.7 Flash on every displayed benchmark row, even where the change is closer to a routine update than a leap.
This is what makes the three-week interval unusual. Google did not replace a weak model with a competitive one. It replaced a competitive model with one that now clears Terra on both the company's table and an independent composite index.
The Price-to-Performance Ratio Is the Product
Gemini 3.8 Flash keeps the introductory 3.7 Flash price through December 31, 2026: USD 0.75 per million input tokens and USD 3.75 per million output tokens. GPT-5.6 Terra is priced at USD 2 input and USD 12 output.
That makes Gemini's input rate 62.5% lower and its output rate 68.75% lower. Put differently, Terra costs about 2.7 times as much for input and 3.2 times as much for output before caching, tool charges, or different token usage enter the calculation.
Price-to-performance is not one number. It is closer to comparing how far two cars travel on a litre of fuel while also checking which one reaches the destination first. A cheap token is not useful if the model needs three attempts, and a high benchmark score is not economical if the model spends several times as many reasoning tokens reaching it.
Artificial Analysis supplies the more useful completed-task view. Its High-effort Gemini 3.8 Flash configuration scores 59 and costs a weighted average of USD 0.58 per Intelligence Index task. Medium scores 57 at USD 0.41, while Low scores 52 at USD 0.24. Terra Low has the cheapest absolute task cost in the release comparison at USD 0.10, but it scores 41. Gemini is not the minimum possible bill. It is buying substantially more measured capability while remaining inexpensive.
Terra's strongest Max setting scores 57 on the same index, two points behind Gemini High. In Artificial Analysis's release comparison, Terra Max produces output at roughly 101 tokens per second while Gemini High reaches about 305. Gemini is cheaper by list price, scores higher at the top tested configuration, and generates output roughly three times as fast.
The introductory discount expires on January 1, 2027. Google's regular price will become USD 1.50 input and USD 7.50 output. Even after the increase, Gemini remains cheaper than Terra's current USD 2 and USD 12 rates. The margin simply stops being as dramatic.
Google Made the Model Work Harder
The strongest caveat comes from Google itself. Gemini 3.8 Flash improves difficult tasks partly by executing more reasoning steps and calling tools repeatedly. At higher effort, it may use more tokens to maximize performance.
That means the per-token price stayed flat while the number of tokens required by a task may not have. Google explicitly recommends lower effort settings or Gemini 3.7 Flash for applications where compute efficiency is the primary constraint. The old model remains supported because 3.8 is not automatically the cheapest choice for every workload.
This distinction matters for agentic tasks. A stronger model can reduce failed loops, repair its own mistakes, and finish jobs a cheaper model abandons. It can also keep reasoning after the useful work is already done. The only reliable production metric is cost per accepted result, measured with the same tools, permissions, and stopping rules a real application uses.
The independent numbers are encouraging because they already include the model's actual token use. Gemini High's USD 0.58 average task cost is higher than Medium's USD 0.41, but the score also rises from 57 to 59. Developers can choose where on that curve their task belongs instead of paying the maximum reasoning bill for every request.
Flash Cyber Is a Different Product
Gemini 3.8 Flash Cyber shares the public model's underlying intelligence but is not generally available. Google is giving prioritized access to government authorities, critical-infrastructure operators, and software maintainers through the Fairwind Program.
Google reports 47.2% pass@1 on CWE-Bench, close to a leading frontier model at 47.8%, and says an internal vulnerability-discovery benchmark across 20 programming languages exceeded 70%. Chrome's security team reported 2.6 times more correct patches than larger commercial models, while Wiz reported higher recall at lower cost on its internal penetration-testing benchmark.
Those results are relevant to how 3.8 Flash was trained, but they should not be attributed directly to the public model. Flash ships with stronger restrictions against cyber offense and chemical, biological, radiological, and nuclear misuse. Flash Cyber has more permissive mitigations for approved defensive work.
Google also says the 3.8 family improved resistance to prompt injection, citing testing by Gray Swan. The public launch post does not provide enough detail to turn that statement into a general security guarantee. A model that resists one benchmark can still be manipulated through another tool, document, or application boundary.
The Actual Story
Gemini 3.8 Flash is the release 3.6 Flash was supposed to foreshadow. It is a fast model with a Flash rate card that now beats GPT-5.6 Terra on 13 of Google's 14 directly comparable benchmark rows and edges Terra's best configuration on an independent intelligence index.
It is not uniformly better than the frontier. Opus 5 and GPT-5.6 Sol still win important rows. Google's table is vendor-reported. Higher performance comes partly from spending more reasoning and tool calls, which can narrow the difference between token price and task price.
None of that breaks the central result. Three weeks after 3.7 Flash made Google competitive with Terra, 3.8 Flash moves ahead while keeping the same introductory rate. The strongest version scores higher than Terra Max, outputs more than twice as fast, and charges a fraction of Terra's token price.
Gemini 3.7 Flash was the cheap alternative. Gemini 3.8 Flash is the cheaper leader.
Previously on TheQuery: