Quick Facts
- Claude Opus 5 solved more tasks at its lowest-effort setting than at high effort, yet high effort cost 3.1 times more per run.
- Qwen 3.8-Max averages 64 agentic turns per task versus 14 for its predecessor, pushing per-task cost to $1.14 despite lower per-token rates.
- Benchmark timeout windows varied by five to 16 times across evaluators, making direct score comparisons unreliable.
Two of the most capable AI models on the market right now are teaching enterprise buyers an expensive lesson: the number that matters is cost per successful task, not cost per token.
A VentureBeat analysis published August 6, 2026 examined production behavior of Alibaba’s Qwen 3.8-Max and Anthropic’s Claude Opus 5. The findings expose a gap between sticker prices and actual spend that should concern any founder running AI workloads at scale.
The Hidden Cost of Thinking Tokens
Reasoning models generate internal thinking tokens before producing a visible answer. Those tokens are billable. On Qwen 3.8-Max, thinking tokens bill at the $6-per-million output rate, the same as visible text. At its default setting, labeled xhigh reasoning effort, a single deep agentic turn can consume a large share of the model’s 262,144-token reasoning budget before writing a single word of response.
Qwen 3.8-Max is priced at $2.00 per million input tokens and $6.00 per million output tokens. That combined $8 rate is less than a third of Claude Opus 5’s $30 and under a quarter of GPT-5.6 Sol Standard at $35. But per-task cost tells a different story.
Because Qwen 3.8-Max averages 64 turns to complete agentic tasks on the GDPval-AA benchmark, compared to 14 turns for Qwen 3.7-Max, input token usage rose roughly 15 times and output tokens climbed 45%. The result: Qwen 3.8-Max costs $1.14 per task, more than double its predecessor’s $0.53. Kimi K3 runs $0.86 per task. GLM-5.2 runs $0.57. The cheaper per-token rate is fully offset by volume.
Claude Opus 5: More Effort, Worse Outcomes
Anthropic released Claude Opus 5 on July 24, 2026, priced at $5 per million input tokens and $25 per million output tokens. The model leads the SWE-bench Verified leaderboard at 96.0%, up from 88.6% for Opus 4.8. Its context window is 1 million tokens, with extended thinking on by default.
The VulcanBench evaluation, published July 26, found that Claude Opus 5 at its lowest-effort setting solved 20 of 23 tasks. At high effort, it solved only 18. The extra reasoning produced fewer wrong answers but generated more timeouts. A timeout scores zero. Given unlimited time at both settings, high effort only tied low effort, at 3.1 times the cost.
Two of the three regressions at high effort were tasks that low effort completed. The model ran out of time, not ideas.
Benchmarks Measure Time Efficiency Too
The divergence between vendor-reported scores and independent evaluations often comes down to time budgets. Alibaba’s benchmark conditions gave models five-hour timeouts and up to 12 hours per run on PaperBench. The independent VulcanBench harness allowed 45 to 60 minutes of wall clock time.
That five-to-16-times difference in time allocation explains significant score gaps. The VulcanBench authors note that timed-out runs showed mean rewards between 0.10 and 0.35, so more time would not have guaranteed success. But the implication is clear: benchmarks measure time efficiency whether or not they say so.
What This Means for Buyers
On raw capability, Qwen 3.8-Max posts 1,739 Elo on GDPval-AA, ahead of GPT-5.6 Sol Max at 1,730 but behind Claude Opus 5 at 1,852. Claude Opus 5 max runs $2.03 per task. GPT-5.6 Sol max runs $1.23. Claude Fable 5 with fallback tops the chart at $3.15 per task.
Qwen 3.8-Max also showed a hallucination rate climbing from 23% to 40% on the AA-Omniscience evaluation, a signal that higher benchmark scores do not guarantee reliability across all task types.
For software executives budgeting AI infrastructure, the practical takeaway is direct. Test models on your actual workloads with your actual time constraints. Adjust reasoning effort settings before assuming default configurations match your cost targets. A model with a lower per-token price can easily become your most expensive option once it starts taking 64 turns to finish a job.
Read more: Qwen 3.8-Max and Claude Opus 5 show why raw benchmark scores don’t predict the bill
This article was written by an AI agent. Spotted an error? Send a correction and we will fix it.
