Quick Facts
- Qwen3.8-Max scores 86.1 on OSWorld-Verified, ahead of GPT-5.6 Sol Max (83.2) and Fable 5 (85.0)
- The model has 2.4 trillion total parameters but activates only 95 billion per request to reduce inference costs
- Alibaba priced the model at $2.00 per 1M input tokens and $6.00 per 1M output tokens via its API
Alibaba’s Qwen team released Qwen3.8-Max on August 3, a 2.4-trillion-parameter mixture-of-experts model targeting autonomous software engineering and long-horizon enterprise tasks. The company previewed the model at the World AI Conference in Shanghai on July 19 before making it fully available this week.
On OSWorld-Verified, a benchmark testing computer-use agents interacting with desktop environments, Qwen3.8-Max posted a score of 86.1. That beat GPT-5.6 Sol Max at 83.2, Fable 5 at 85.0, and Gemini 3.1 Pro at 76.2.
The gains against its own predecessor are large. On DeepSWE 1.1, the model moved from 21.6 to 56.6. On FrontierSWE, it went from 40.7 to 73.5. On JobBench, it climbed from 31.3 to 53.4.
Performance is not uniform across all benchmarks. Fable 5 still leads on FrontierSWE with a score of 88.8, compared to Qwen3.8-Max’s 73.5. On professional workflow tasks, Fable 5 also edges ahead, scoring 75.9 on CoWorkBench versus Qwen3.8-Max’s 74.8, and 57.4 on JobBench versus 53.4.
On PaperBench, which tests whether a model can reproduce research paper results end to end, Qwen3.8-Max scored 93.0. That placed it ahead of GPT-5.6 Sol at 90.5, Fable 5 at 88.8, and Claude Opus 4.8 at 80.3.
Third-party ranking service Arena placed Qwen3.8-Max fifth on its Text Arena leaderboard and second on Vision Arena. On the Frontend Code Arena leaderboard, it ranked fourth with a score of 1,668, behind Claude Opus 5 Max at 1,705 and Kimi K3 Max at 1,676.
The model supports a context window of up to 983,616 tokens with a maximum output of 131,072 tokens. It processes text, images, and video. Thinking is always enabled, with low, high, and xhigh reasoning settings; xhigh is the default.
Speed is a noted limitation. According to Artificial Analysis, Qwen3.8-Max generates output at 46.5 tokens per second through Alibaba’s API. The median for comparable reasoning models sits at 72 tokens per second.
Alibaba said the model can autonomously complete software projects spanning more than 10 days, reproduce research papers involving thousands of lines of code, and perform chip-design optimization. The company is positioning the model as an autonomous worker rather than a conversational assistant.
On pricing, the $6.00 per million output tokens undercuts the peer group median of $10.00. Input tokens at $2.00 per million sit slightly above the median of $1.75.
Alibaba said Qwen3.8-Max will be the first Max-class Qwen model it open-sources, reversing a move toward proprietary releases the company made earlier this year. The benchmark scores are vendor-reported and draw from multiple testing harnesses, so direct comparisons across models carry caveats.
