>_TheQuery
// Reading nowStart
← All Articles

Grok 4.6 Ties GPT-5.6 Sol. Where It Falls Short

By Addy · August 12, 2026 · Editorial standards

SpaceXAI, xAI's parent branding since its consolidation into Musk's broader stack, shipped Grok 4.6 on August 12. It's a genuinely strong release, with the important caveat that "genuinely strong" and "ahead of the field" turn out to be different claims once you read past the launch post's prose and into its own published table.

This isn't a bigger base model. xAI didn't publish a parameter count, and the official description is explicit that this is a post-training upgrade on Grok 4.5, a longer supplemental training run on curated reasoning and engineering data, a revised optimizer, regenerated fine-tuning trajectories, and reinforcement learning across coding, kernel optimization, web development, computer-aided design, and agent environments. It ships with a 500,000-token context window, text and image input with text-only output, a February 1, 2026 knowledge cutoff, and a new xhigh reasoning-effort tier alongside the existing low, medium, and high settings. It's available now in Cursor, xAI's own Grok Build, the API, OpenRouter, Vercel, and Cloudflare, with Cursor and Grok Build offering double included usage for the first week. The pitch is long-running agentic workflows specifically, not a raw intelligence jump, and on that specific pitch, the independently verified numbers actually back it up. On several others, they don't.

The Efficiency Win Is Real

Artificial Analysis evaluated Grok 4.6 directly rather than just relaying xAI's own claims, and the standout result isn't a leaderboard rank, it's how little Grok 4.6 needed to get there. On AA-Briefcase, Artificial Analysis's own private benchmark of long-horizon agentic knowledge work, the kind of multi-step research and drafting work that doesn't reduce cleanly to a pass or fail test, Grok 4.6 resolves tasks in roughly 53 turns and 0.5 billion input tokens on average. Claude Opus 5, working through comparable tasks at a similar quality tier, needs roughly 103 turns and 2.0 billion input tokens to get there. That's about a quarter of the token spend for work landing in the same tier on rubric grading, presentation, and analytical quality, not one strong dimension propping up a weak score elsewhere.

Grok 4.6 also leads outright on two individual benchmarks: GDPVal-AA v2, a knowledge-work benchmark, at 1753 against Fable 5 Max's 1741 and GPT-5.6 Sol Max's 1728, and Harvey LAB, a legal-reasoning benchmark, where it comes out ahead of both rivals as well. Those are genuine, independently confirmed first-place finishes, not rounding-error ties.

The Official Table

It's worth showing the actual table rather than summarizing it, since the pattern only really lands when you can see all ten rows next to each other the way xAI's own page presents them. This is reproduced directly from x.ai's launch page, with xAI's own note preserved: third-party scores are the best of self-reported or publicly available results, and a dash means no score was available.

BenchmarkGrok 4.6 HighGrok 4.5 HighGPT-5.6 Sol MaxFable 5 Max
AA Intelligence Index61566162
GDPVal-AA v21753152617281741
CursorBench v3.269.9%66.7%67.2%70.5%
DeepSWE v1.165.9%54%73%70%
FrontierCode v1.1 (Extended)61.3%56.6%60.6%63.6%
APEX-Agents57.5%47.1%56.7%59.2%
Terminal-Bench v3.026%15.7%34.6%34.1%
APEX-SWE56.4%53.6%not reported58.8%
AA-Briefcase1577131315021574
Harvey LAB (Vals)15.8%12.9%2.5%11.3%

Counted row by row, Fable 5 Max posts the best published score on five of the ten evaluations. Grok 4.6 leads on three: GDPVal-AA v2, AA-Briefcase, and Harvey LAB. GPT-5.6 Sol Max leads on two: DeepSWE v1.1 and Terminal-Bench v3.0, the same two rows where coding and terminal work specifically are being tested. Grok 4.5, last generation's model, doesn't lead a single row, which is the correct and expected outcome and the clearest confirmation that the comparison itself is being read fairly rather than stacked.

That row count is also the cleanest possible evidence for the point about composite scores from earlier: Grok 4.6 ties for the second-best average across nine blended benchmarks while winning the outright top score on only three of the ten rows shown. Both facts are true about the same model at the same time.

The generational jump against Grok 4.5 is unambiguous in one direction, though. Every score in that table moved up, and by a wide margin in most cases: AA Intelligence Index up 5 points, AA-Briefcase up 264 Elo, Terminal-Bench v3.0 up 10.3 points, DeepSWE up 11.9 points, APEX-Agents up 10.4 points. That's a real, substantial step up from the previous version in five weeks. It just isn't the same claim as leading the field outright, and xAI's own launch prose leans on the first claim while mostly staying quiet about where the second one doesn't hold.

What Tying on the Index Actually Means

A composite score works the way a report card GPA does.

Two students can post the same GPA while one is consistently solid across every subject and the other is excellent in half their classes and mediocre in the rest. That's exactly what the table above shows: Grok 4.6 ties GPT-5.6 Sol Max on the nine-benchmark average while actually winning fewer individual rows than Fable 5 Max, which sits one point ahead on the average and five rows ahead on outright wins. The average and the win count are both accurate. They're just answering different questions.

The Row Its Own Post Doesn't Mention

The clearest individual gap is Terminal-Bench v3.0, the newer, harder benchmark built specifically to stay unsaturated at the top of the field, the same one we covered as still too new for most labs to have widely adopted. Grok 4.6 scores 26% there, against 34.6% for GPT-5.6 Sol Max and 34.1% for Fable 5 Max, last of the four models on xAI's own table, a meaningful 8.6-point gap to the leader.

What's actually notable isn't the score, it's where it appears. MarkTechPost's read of the launch materials points out directly that this specific row never comes up in the launch post's prose. It's published plainly in the table underneath, available to anyone who scrolls past the summary, just never mentioned in the part most people actually read. DeepSWE v1.1 follows the same pattern, 65.9% against GPT-5.6 Sol Max's 73%, a real improvement over Grok 4.5 and still the second weakest result on the table. Both of Grok 4.6's clearest losses are coding and terminal work specifically, the categories engineering teams are most likely to actually care about when picking a model for real work rather than knowledge tasks.

The Pricing Has a Cliff at 200,000 Tokens

The headline inference rate is straightforward: $2 per million input tokens, $6 per million output, roughly 60% below GPT-5.6 Sol's $5 and $30. What the headline number leaves out is that it only applies below a 200,000-token prompt. Cross that line and the rate doubles to $4 input and $12 output, and critically, xAI's own documentation states the higher rate applies to every token in that request, not just the tokens past the threshold. A prompt that lands at 210,000 tokens doesn't get charged the premium rate for 10,000 tokens. The whole request gets billed at the higher rate.

Two more details worth knowing before shipping anything on this. The launch page references a faster variant available at double the price, with no separate model ID published anywhere, meaning there's currently no way to reference or test it directly outside whatever xAI exposes through routing. And without a properly set prompt_cache_key or x-grok-conv-id header, requests scatter across different servers and cache hits stop working reliably, meaning the full, uncached price applies by default rather than as an edge case someone might stumble into occasionally.

Where This Sits in a Very Crowded Month

Grok 4.6 is the latest entry in what's been an unusually dense stretch of launches, DeepSeek V4 Pro's move to general availability with an actual independent Artificial Analysis score behind it, Qwen 3.8 Max still waiting on its promised open weights and any independent verification at all, and now this. Measured against that immediate backdrop, Grok 4.6's strongest claim, the efficiency profile, holds up about as well as anything else that's shipped this month, because it's Artificial Analysis's own number, not xAI's. That's a meaningfully higher bar cleared than Qwen managed, and it's worth crediting plainly.

What keeps this from being a clean win is the same discipline this publication has applied to every other launch this month: read the whole table, not the summary. Grok 4.6 is a real, independently verified step forward, genuinely more efficient than the model it's being compared against, genuinely ahead outright on two of ten categories. It is also a model whose own launch table shows exactly where that isn't true, in a benchmark its own prose never brings up, priced in a way that punishes exactly the kind of long-context work it's being marketed for the moment you cross one specific, easy-to-miss line.

Previously on TheQuery:

Sources

  1. Introducing Grok 4.6
  2. Grok 4.6 Returns SpaceXAI to the Intelligence Frontier and Leads on Cost Efficiency
  3. SpaceXAI Releases Grok 4.6: A 500K-Context Frontier Model Tuned for Long-Running Agents, Coding, and Knowledge Work