>_TheQuery
// Reading nowStart
← All Articles

Qwen 3.8 Max's Own Benchmark Chart Leaves Out Kimi K3 and Claude Opus 5. It Still Doesn't Clearly Win.

By Addy · August 3, 2026

Alibaba shipped Qwen 3.8 Max today with a sixteen-panel benchmark chart comparing it against six models: its own predecessors Qwen 3.7 Max and Qwen 3.7 Plus, plus Claude Opus 4.8, Claude Fable 5, Gemini 3.1 Pro, and GPT-5.6 Sol. Two names anyone following this space would expect to see are missing entirely: Kimi K3, and Claude Opus 5.

Neither omission looks accidental once you know the dates.

The Two Names That Should Have Been There

Kimi K3 is the obvious size and category peer here. It's an open-weight model at a comparable scale, 2.8 trillion parameters against Qwen 3.8 Max's 2.4 trillion, and it currently holds the top spot among open-weight models on the independent Artificial Analysis Intelligence Index. If Alibaba wanted to show Qwen 3.8 Max leading the open-weight field, the model actually leading that field right now was the one comparison that would have tested that claim directly. It isn't in the chart.

Claude Opus 5 is the sharper omission. It shipped July 24, eleven days before this chart went out. By the time Alibaba published these sixteen panels, Opus 5 had already overtaken Fable 5 on the independent Intelligence Index, 61 to 60, the model actually holding the top overall spot in the industry right now. Alibaba compared against Opus 4.8 instead, a model Opus 5 had already replaced weeks earlier. That's not comparing against an old result because a new one wasn't out yet. The new one had been out for over a week.

Where Qwen 3.8 Max Falls Short

Here's the part that matters more than the omissions: even with the two most inconvenient comparisons removed, the chart Alibaba actually published doesn't clearly support “second only to Fable 5” either.

On the benchmarks Alibaba itself frames as central to this model, coding and agentic workflows, Fable 5 leads Qwen 3.8 Max repeatedly: SWE-Pro (80.0 to 67.7), FrontierSWE (88.8 to 73.5), QwenReactBench (1770 to 1724), CoWorkBench (75.9 to 74.8), JobBench (57.4 to 53.4), Vision2Web (70.5 to 69.0), and MobileWorld (85.5 to 77.8). GPT-5.6 Sol beats it too, on TerminalBench-2.1 (88.8 to 86.6) and Agents' Last Exam (53.6 to 52.4). That's nine panels, out of sixteen, where a model still on the chart outperforms Qwen 3.8 Max, on exactly the categories, software engineering and long-horizon AI agent work, that Alibaba's own announcement describes as the model's target use case.

Where Does Qwen 3.8 Max Actually Lead

Where Qwen 3.8 Max does lead clearly is a different cluster: PaperBench, BabyVision, CharXiv, ERQA, PerceptionBench, LVBench, and OSWorld-Verified. Look at what those have in common. They're vision, perception, and research-reproduction tasks, not the coding and agentic categories the launch is actually being sold on. Alibaba's own chart, read plainly, shows a model that's genuinely strong at multimodal perception and comparatively behind on the exact work its own positioning leads with.

Alibaba's release notes include a real caveat worth repeating rather than skipping past: competing models were evaluated in their own preferred harnesses, Claude models through Claude Code, GPT-5.6 Sol through Codex, which the release itself describes as making cross-lab comparisons directional rather than exact. That caveat cuts both ways. It's a legitimate reason to treat any single bar with some skepticism. It is not a reason the two most current, most relevant comparison points were left off the chart entirely.

Leave out the field's actual open-weight leader and the model that already holds the top overall spot, and a chart can say almost anything. Include everything that was already public the day this one was built, and it says something closer to: strong at vision, competitive at best on coding, and not actually ahead of the field it's being measured against.

Sources:

Previously on TheQuery: