A task with its own rules.
A visual mathematics benchmark that tests whether a model can solve math problems grounded in diagrams, equations, figures, and other visual inputs.
- Organisation
- Qwen
- Version
- 2026
Multimodal / Benchmark profile
A visual mathematics benchmark that tests whether a model can solve math problems grounded in diagrams, equations, figures, and other visual inputs.
Observed results
13 ranked models · higher is better · labels show rank and score
| Model | Rank | Model ID | Provider | Score | Unit |
|---|---|---|---|---|---|
| Qwen3.8-Flash-Next | #1 | qwen-3-8-flash-next | Alibaba Cloud | 95.7 | percent |
| Qwen3.8-27B | #2 | qwen-3-8-27b | Alibaba Cloud | 94.6 | percent |
| Kimi K3 | #3 | kimi-k3 | Moonshot AI | 94.3 | percent |
| Qwen3.7-Plus | #4 | qwen-3-7-plus | Alibaba Cloud | 90.3 | percent |
| Qwen3.5 397B A17B | #5 | qwen3-5-397b | Alibaba Cloud | 88.6 | percent |
| Qwen3.6 Plus | #6 | qwen3-6-plus | Alibaba Cloud | 88 | percent |
| Kimi K2.6 | #7 | kimi-k2-6 | Moonshot AI | 87.4 | percent |
| Qwen3.5-122B-A10B | #8 | qwen3-5-122b-a10b | Alibaba Cloud | 86.2 | percent |
| Qwen3.5 27B | #9 | qwen3-5-27b | Alibaba Cloud | 86 | percent |
| Qwen3.5-35B-A3B | #10 | qwen3-5-35b-a3b | Alibaba Cloud | 83.9 | percent |
| GPT-5.2 | #11 | gpt-5-2 | OpenAI | 83 | percent |
| Gemma 4 12B Unified | #12 | gemma-4-12b | 79.7 | percent | |
| Claude Opus 4.5 | #13 | claude-opus-4-5 | Anthropic | 74.3 | percent |
One best compatible score per canonical product · higher is better
| Rank | Model | Weights | Evidence | Score |
|---|---|---|---|---|
| 01 | Qwen3.8-Flash-NextAlibaba Cloud | Open weights | Public reference | 95.7% |
| 02 | Qwen3.8-27BAlibaba Cloud | Open weights | Public reference | 94.6% |
| 03 | Kimi K3Moonshot AI | Open weights | Source carrier | 94.3% |
| 04 | Qwen3.7-PlusAlibaba Cloud | Closed weights | Source carrier | 90.3% |
| 05 | Qwen3.5 397B A17BAlibaba Cloud | Open weights | Source carrier | 88.6% |
| 06 | Qwen3.6 PlusAlibaba Cloud | Closed weights | Source carrier | 88% |
| 07 | Kimi K2.6Moonshot AI | Open weights | Source carrier | 87.4% |
| 08 | Qwen3.5-122B-A10BAlibaba Cloud | Open weights | Source carrier | 86.2% |
| 09 | Qwen3.5 27BAlibaba Cloud | Open weights | Source carrier | 86% |
| 10 | Qwen3.5-35B-A3BAlibaba Cloud | Open weights | Source carrier | 83.9% |
Each distinct published measurement is retained. Same-snapshot canonical and configuration projections appear once.
| Published model / configuration | Score | Evidence & protocol | Source & dates |
|---|---|---|---|
Qwen3.8-Flash-Next (xhigh default thinking configuration)Qwen3.8-Flash-NextExact identityqwen-3-8-flash-next-xhigh Canonical product: qwen-3-8-flash-next | 95.7% | Public referencesource-checkedVersion & systemWith CI system:qwen3-8-flash-next:mathvision-with-ci | Qwen3.8-Flash-Next current official launch pageObserved Checked |
Qwen3.8-27B (xhigh default thinking configuration) as published in Qwen's Qwen3.8-Flash-Next comparison tableQwen3.8-27BExact identityqwen-3-8-27b-xhigh Canonical product: qwen-3-8-27b | 94.6% | Relative comparisonprovider-reportedVersion & systemWith CI system:qwen3-8-flash-next-comparison:cell:vision:mathvision-with-ci:1 | Qwen3.8-Flash-Next launch and official provider evaluationsObserved Checked |
Qwen3.8-Flash-Next (xhigh default thinking configuration)Qwen3.8-Flash-NextExact identityqwen-3-8-flash-next-xhigh Canonical product: qwen-3-8-flash-next | 90.6% | Public referencesource-checkedVersion & systemWithout CI system:qwen3-8-flash-next:mathvision-no-ci | Qwen3.8-Flash-Next current official launch pageObserved Checked |
Qwen3.7-Plus (provider-published configuration) as published in Qwen's Qwen3.8-Flash-Next comparison tableQwen3.7-PlusExact identityqwen-3-7-plus-unspecified Canonical product: qwen-3-7-plus | 90.3% | Relative comparisonprovider-reportedVersion & systemWithout CI system:qwen3-8-flash-next-comparison:cell:vision:mathvision-no-ci:2 | Qwen3.8-Flash-Next launch and official provider evaluationsObserved Checked |
Qwen3.8-27B (xhigh default thinking configuration) as published in Qwen's Qwen3.8-Flash-Next comparison tableQwen3.8-27BExact identityqwen-3-8-27b-xhigh Canonical product: qwen-3-8-27b | 90% | Relative comparisonprovider-reportedVersion & systemWithout CI system:qwen3-8-flash-next-comparison:cell:vision:mathvision-no-ci:1 | Qwen3.8-Flash-Next launch and official provider evaluationsObserved Checked |
Qwen3.7-Plus (provider-published configuration) as published in Qwen's Qwen3.8-Flash-Next comparison tableQwen3.7-PlusExact identityqwen-3-7-plus-unspecified Canonical product: qwen-3-7-plus | 88.7% | Relative comparisonprovider-reportedVersion & systemWith CI system:qwen3-8-flash-next-comparison:cell:vision:mathvision-with-ci:2 | Qwen3.8-Flash-Next launch and official provider evaluationsObserved Checked |
Claude Opus 4.6 (Max) as published in Qwen's Qwen3.8-Flash-Next comparison tableClaude Opus 4.6Exact identityclaude-opus-4-6-max Canonical product: claude-opus-4-6 | 65.5% | Relative comparisonprovider-reportedVersion & systemWithout CI system:qwen3-8-flash-next-comparison:cell:vision:mathvision-no-ci:3 | Qwen3.8-Flash-Next launch and official provider evaluationsObserved Checked |
Qwen3.8-27B (xhigh)Qwen3.8-27BExact identityqwen-3-8-27b-xhigh Canonical product: qwen-3-8-27b | 94.6% | Public referenceprovider-reportedVersion & systemWith CI Qwen3.8-27B official VL table | Qwen3.8-27B official model cardObserved Checked |
Qwen3.7-Plus as published by QwenQwen3.7-PlusExact identityqwen-3-7-plus-unspecified Canonical product: qwen-3-7-plus | 90.3% | Relative comparisonprovider-reportedVersion & systemWithout CI Qwen3.8-27B official VL table | Qwen3.8-27B official model cardObserved Checked |
Qwen3.8-27B (xhigh)Qwen3.8-27BExact identityqwen-3-8-27b-xhigh Canonical product: qwen-3-8-27b | 90% | Public referenceprovider-reportedVersion & systemWithout CI Qwen3.8-27B official VL table | Qwen3.8-27B official model cardObserved Checked |
From result to context
A visual mathematics benchmark that tests whether a model can solve math problems grounded in diagrams, equations, figures, and other visual inputs.
The summary shows one best compatible source result per canonical product. Versions, effort settings and execution systems remain attached to the underlying records.
Scores from different versions or harnesses may not be interchangeable. The published source rows preserve those distinctions and their original units. A source result is not automatically an input to a current index.
This catalogue entry preserves available source evidence. Current indices admit only their specifically reviewed tracks and configurations.
Contamination risk: Unknown. Lifecycle: Active.