Alibaba's Qwen3.8-Max Claims to Beat GPT-5.6
Alibaba's Qwen team has unveiled Qwen3.8-Max, a 2.4-trillion-parameter mixture-of-experts multimodal model targeting autonomous software engineering and enterprise AI tasks. The company claims it outperforms leading proprietary models on key benchmarks.
Qwen3.8-Max scored 86.1 on the OSWorld-Verified benchmark, surpassing GPT-5.6 Sol Max at 83.2 and Fable 5 at 85.0. It also posted the highest reported score on OpenAI's PaperBench, which tests how well AI agents can reconstruct scientific research from experimental data. Independent verification of these results is still pending.
Alibaba closes the agentic benchmark gap with American frontier labs faster than the export controls intended.
