Reasoning
Mathematical and scientific problem solving.
Score and research range
- Full-precision point
- 112.04024297326077
- Range
- 109.42–114.66 · 90% Research Range
- Scope
- Conditional fuller R0 A
Score includes qualified prediction of missing evidence.
Model profile / Alibaba Cloud
Qwen's open-weight 27B dense vision-language model with native 262,144-token context, extensible to 1,000,000 tokens, and configurable xhigh, medium and low reasoning effort.
Four areas of evidence
Mathematical and scientific problem solving.
Score includes qualified prediction of missing evidence.
Software engineering and programming.
Score uses the admitted direct native evidence under the selected method.
Factual reliability and instruction following.
Score uses the admitted direct native evidence under the selected method.
Tool use, workflows and professional tasks.
Score uses the admitted direct native evidence under the selected method.
Current index snapshot · 2026-09-16. Each category uses its own scale; scores across categories are not directly comparable. Evidence counts describe the retained Capabilities inputs; a family can contain several tasks or configurations. A point lead does not establish superiority.
Performance in context
Qwen3.8-27B is highlighted wherever a compatible measurement is available.
Qwen3.8-27B: Regional/tier or mutable endpoint-to-checkpoint binding unresolved; no price inherited
Lumina Capabilities Index
The current Capabilities cohort, with this model highlighted.
Chart loads as you explore
Qwen3.8-27B: No exact base-profile release and documented 10K output measurement; other effort/protocol/alias records not substituted
Lumina Capabilities Index
The current Capabilities cohort, with this model highlighted.
Chart loads as you explore
Explore the detail
Score uses the admitted direct native evidence under the selected method.
Score uses the admitted direct native evidence under the selected method.
Score includes qualified prediction of missing evidence.
Score uses the admitted direct native evidence under the selected method.
Score uses the admitted direct native evidence under the selected method.
| Benchmark | Result | Published configuration | Source |
|---|---|---|---|
| Agents Last ExamPass@1 | 20.4% | Qwen3.8-27B (xhigh)Qwen3.8-27B official text table | Qwen3.8-27B official model card ↗Observed 2026-08-15 · checked 2026-08-15provider-reported · reference-only |
| Agents Last ExamScore | 42.9% | Qwen3.8-27B (xhigh)Qwen3.8-27B official text table | Qwen3.8-27B official model card ↗Observed 2026-08-15 · checked 2026-08-15provider-reported · reference-only |
| AndroidWorld2026 | 81.9% | Qwen3.8-27B (xhigh)Qwen3.8-27B official VL table | Qwen3.8-27B official model card ↗Observed 2026-08-15 · checked 2026-08-15provider-reported · reference-only |
| BabyVisionWith CI | 85.6% | Qwen3.8-27B (xhigh)Qwen3.8-27B official VL table | Qwen3.8-27B official model card ↗Observed 2026-08-15 · checked 2026-08-15provider-reported · reference-only |
| BabyVisionWithout CI | 65.7% | Qwen3.8-27B (xhigh)Qwen3.8-27B official VL table | Qwen3.8-27B official model card ↗Observed 2026-08-15 · checked 2026-08-15provider-reported · reference-only |
| Claw-EvalAverage | 56.9% | Qwen3.8-27B (xhigh)Qwen3.8-27B official VL table | Qwen3.8-27B official model card ↗Observed 2026-08-15 · checked 2026-08-15provider-reported · reference-only |
| Claw-EvalPass@3 | 57.4% | Qwen3.8-27B (xhigh)Qwen3.8-27B official VL table | Qwen3.8-27B official model card ↗Observed 2026-08-15 · checked 2026-08-15provider-reported · reference-only |
| ERQA2026 | 65.5% | Qwen3.8-27B (xhigh)Qwen3.8-27B official VL table | Qwen3.8-27B official model card ↗Observed 2026-08-15 · checked 2026-08-15provider-reported · reference-only |
| JobBench2026 | 33.4% | Qwen3.8-27B (xhigh)Qwen3.8-27B official text table | Qwen3.8-27B official model card ↗Observed 2026-08-15 · checked 2026-08-15provider-reported · reference-only |
| MathVisionWith CI | 94.6% | Qwen3.8-27B (xhigh)Qwen3.8-27B official VL table | Qwen3.8-27B official model card ↗Observed 2026-08-15 · checked 2026-08-15provider-reported · reference-only |