Reasoning
Mathematical and scientific problem solving.
Model profile / Alibaba Cloud
Qwen's managed 2.4T-parameter flagship API has a 1M-token context window and multimodal features. Qwen also publishes the related text checkpoint weights under the Qwen3.8 Max License; managed-only features remain endpoint-specific.
Four areas of evidence
Mathematical and scientific problem solving.
Software engineering and programming.
Factual reliability and instruction following.
Tool use, workflows and professional tasks.
Current index snapshot · 2026-09-16. Each category uses its own scale; scores across categories are not directly comparable. Counts describe published catalogue evidence. A point lead does not establish superiority.
Performance in context
Qwen3.8 Max has no current Capabilities rating. Explore the rated models below.
Lumina Capabilities Index
The current Capabilities cohort, with available operational measurements.
Chart loads as you explore
Lumina Capabilities Index
The current Capabilities cohort, with available operational measurements.
Chart loads as you explore
Explore the detail
No current index estimate. Available benchmark measurements and source records are retained below.
| Benchmark | Result | Published configuration | Source |
|---|---|---|---|
| FrontierMath v2 Tiers 1-3v2 / task implementation 2.0.0 | 74.737% | Qwen3.8 Max (xhigh)Epoch AI benchmark runner | Epoch AI permanent refresh source ↗Observed 2026-08-04 · checked 2026-08-17source-checked · excluded |
| FrontierMath v2 Tier 4v2 / task implementation 2.0.0 | 46.341% | Qwen3.8 Max (xhigh)Epoch AI benchmark runner | Epoch AI permanent refresh source ↗Observed 2026-08-04 · checked 2026-08-17source-checked · excluded |
| GPQA Diamond1.0.11 | 92.677% | Qwen3.8 Max (xhigh)Epoch AI benchmark runner | Epoch AI permanent refresh source ↗Observed 2026-08-04 · checked 2026-08-17source-checked · ranking-eligible |
| SimpleQA Verified1.2.0 | 45.8% | Qwen3.8 Max (xhigh)Epoch AI benchmark runner | Epoch AI permanent refresh source ↗Observed 2026-08-27 · checked 2026-09-01independently-verified · reference-only |
| AA AutomationBench1.0.6 | 39.8% | Qwen3.8 Max as published by Z.AIZ.AI GLM-5.3 official comparison table | GLM-5.3: Frontier Coding with Emergent Cyber Capabilities ↗Observed 2026-08-14 · checked 2026-08-15provider-reported · excluded |
| Agents Last ExamALE-CLI | 27% | Qwen3.8 Max as published by Z.AIZ.AI GLM-5.3 official comparison table | GLM-5.3: Frontier Coding with Emergent Cyber Capabilities ↗Observed 2026-08-14 · checked 2026-08-15provider-reported · excluded |
| Humanity's Last Exam with toolswith tools | 56.2% | Qwen3.8 Max as published by Z.AIZ.AI GLM-5.3 official comparison table | GLM-5.3: Frontier Coding with Emergent Cyber Capabilities ↗Observed 2026-08-14 · checked 2026-08-15provider-reported · excluded |
| NL2Repo2026 | 55.9% | Qwen3.8 Max as published by Z.AIZ.AI GLM-5.3 official comparison table | GLM-5.3: Frontier Coding with Emergent Cyber Capabilities ↗Observed 2026-08-14 · checked 2026-08-15provider-reported · excluded |
| CyberGymrolling | 78.5% | Qwen3.8 Max as published by Z.AIZ.AI GLM-5.3 official comparison table | GLM-5.3: Frontier Coding with Emergent Cyber Capabilities ↗Observed 2026-08-14 · checked 2026-08-15provider-reported · excluded |
| DeepSWE1.1 | 56.6% | Qwen3.8 Max as published by Z.AIZ.AI GLM-5.3 official comparison table | GLM-5.3: Frontier Coding with Emergent Cyber Capabilities ↗Observed 2026-08-14 · checked 2026-08-15provider-reported · excluded |