Alibaba released Qwen3.8-Max on the morning of August 3, Beijing time, pitching it at coding and what the company calls "Cowork" professional office work. The Chinese primary coverage puts the announcement at 11:00 Beijing time, or 03:00 UTC.
The specification
The model is a mixture of experts with 2.4 trillion total parameters. Chinese-language coverage puts the activated count at 95 billion per request — about four percent — while the English write-ups say Alibaba did not disclose it. Context runs to 1 million tokens, split as 991K input and 131K output. Pricing is $2.00 per million input tokens, $6.00 output, $0.25 cached.
The table everyone quoted
Alibaba published wins on Terminal-Bench 2.1 (86.6), PaperBench (93.0), GPQA Diamond (92.6) and OSWorld-Verified (86.1), plus a run of autonomy demonstrations: a 16-day unsupervised build producing 265 commits, 127 pull requests and 151 issues; a 125-hour paper reproduction that beat the published method by 2.71 points on AIME24; a chip-design run cutting gate count from 8,298 to 678; and a live competition entry finishing ahead of 458 of 526 human teams inside 24 hours. Every one of those is vendor-run and none is independently reproducible.
The rows that were not quoted
On agentic software engineering, the same table is a loss sheet. SWE-bench Pro: 67.7 against Claude Fable 5's 80.0. DeepSWE 1.1: 56.6 against 70.0. FrontierSWE: 73.5 against 88.8. Humanity's Last Exam: 43.6 against 53.3. Even the flagship Terminal-Bench number sits under GPT-5.6 Sol's 88.8. Part of the comparison runs on NL2Repo-Bench, Alibaba's own benchmark, which nobody outside the company can reproduce.
"Open" is a calendar entry
The sovereignty and cost arguments being made for this model rest on weights that do not exist publicly. On August 3 there was no Qwen3.8 repository on Hugging Face at all; the newest items were ASR models updated twelve days earlier.
