A task with its own rules.
A grounded visual reasoning benchmark focused on evidence-based question answering over real images.
- Organisation
- Qwen
- Version
- 2026
Multimodal / Benchmark profile
A grounded visual reasoning benchmark focused on evidence-based question answering over real images.
Observed results
8 ranked models · higher is better · labels show rank and score
| Model | Rank | Model ID | Provider | Score | Unit |
|---|---|---|---|---|---|
| Qwen3.8-Flash-Next | #1 | qwen-3-8-flash-next | Alibaba Cloud | 72.3 | percent |
| Qwen3.7-Plus | #2 | qwen-3-7-plus | Alibaba Cloud | 69.8 | percent |
| Qwen3.8-27B | #3 | qwen-3-8-27b | Alibaba Cloud | 65.5 | percent |
| GPT-5.4 | #4 | gpt-5-4 | OpenAI | 65.4 | percent |
| Muse Spark | #5 | muse-spark | Meta | 64.7 | percent |
| Qwen3.6 27B | #6 | qwen3-6-27b | Alibaba Cloud | 62.5 | percent |
| Grok 4.20 | #7 | grok-4-20 | xAI | 54.1 | percent |
| Claude Opus 4.6 | #8 | claude-opus-4-6 | Anthropic | 51.6 | percent |
One best compatible score per canonical product · higher is better
| Rank | Model | Weights | Evidence | Score |
|---|---|---|---|---|
| 01 | Qwen3.8-Flash-NextAlibaba Cloud | Open weights | Public reference | 72.3% |
| 02 | Qwen3.7-PlusAlibaba Cloud | Closed weights | Source carrier | 69.8% |
| 03 | Qwen3.8-27BAlibaba Cloud | Open weights | Public reference | 65.5% |
| 04 | GPT-5.4OpenAI | Closed weights | Source carrier | 65.4% |
| 05 | Muse SparkMeta | Closed weights | Source carrier | 64.7% |
| 06 | Qwen3.6 27BAlibaba Cloud | Open weights | Source carrier | 62.5% |
| 07 | Grok 4.20xAI | Closed weights | Source carrier | 54.1% |
| 08 | Claude Opus 4.6Anthropic | Closed weights | Source carrier | 51.6% |
Each distinct published measurement is retained. Same-snapshot canonical and configuration projections appear once.
| Published model / configuration | Score | Evidence & protocol | Source & dates |
|---|---|---|---|
Qwen3.8-Flash-Next (xhigh default thinking configuration)Qwen3.8-Flash-NextExact identityqwen-3-8-flash-next-xhigh Canonical product: qwen-3-8-flash-next | 72.3% | Public referencesource-checkedVersion & system2026 system:qwen3-8-flash-next:erqa | Qwen3.8-Flash-Next current official launch pageObserved Checked |
Qwen3.7-Plus (provider-published configuration) as published in Qwen's Qwen3.8-Flash-Next comparison tableQwen3.7-PlusExact identityqwen-3-7-plus-unspecified Canonical product: qwen-3-7-plus | 69.8% | Relative comparisonprovider-reportedVersion & system2026 system:qwen3-8-flash-next-comparison:cell:vision:erqa:2 | Qwen3.8-Flash-Next launch and official provider evaluationsObserved Checked |
Qwen3.8-27B (xhigh default thinking configuration) as published in Qwen's Qwen3.8-Flash-Next comparison tableQwen3.8-27BExact identityqwen-3-8-27b-xhigh Canonical product: qwen-3-8-27b | 65.5% | Relative comparisonprovider-reportedVersion & system2026 system:qwen3-8-flash-next-comparison:cell:vision:erqa:1 | Qwen3.8-Flash-Next launch and official provider evaluationsObserved Checked |
Claude Opus 4.6 (Max) as published in Qwen's Qwen3.8-Flash-Next comparison tableClaude Opus 4.6Exact identityclaude-opus-4-6-max Canonical product: claude-opus-4-6 | 40.8% | Relative comparisonprovider-reportedVersion & system2026 system:qwen3-8-flash-next-comparison:cell:vision:erqa:3 | Qwen3.8-Flash-Next launch and official provider evaluationsObserved Checked |
Qwen3.7-Plus as published by QwenQwen3.7-PlusExact identityqwen-3-7-plus-unspecified Canonical product: qwen-3-7-plus | 69.8% | Relative comparisonprovider-reportedVersion & system2026 Qwen3.8-27B official VL table | Qwen3.8-27B official model cardObserved Checked |
Qwen3.8-27B (xhigh)Qwen3.8-27BExact identityqwen-3-8-27b-xhigh Canonical product: qwen-3-8-27b | 65.5% | Public referenceprovider-reportedVersion & system2026 Qwen3.8-27B official VL table | Qwen3.8-27B official model cardObserved Checked |
Qwen3.6-27B as published by QwenQwen3.6 27BExact identityqwen3-6-27b-default Canonical product: qwen3-6-27b | 62.5% | Relative comparisonprovider-reportedVersion & system2026 Qwen3.8-27B official VL table | Qwen3.8-27B official model cardObserved Checked |
Claude Opus 4.6 Max as published by QwenClaude Opus 4.6Exact identityclaude-opus-4-6-max Canonical product: claude-opus-4-6 | 40.8% | Relative comparisonprovider-reportedVersion & system2026 Qwen3.8-27B official VL table | Qwen3.8-27B official model cardObserved Checked |
Exact BenchLM registry variant Qwen3.7 Plus; bulk export does not retain a complete upstream harness configuration.Qwen3.7-PlusExact identitySource label without a registered configuration ID Canonical product: qwen-3-7-plus | 69.8% | Source carriersource-checkedVersion & system2026 Source-native system | BenchLM public datasets — 2026-08-01Observed Checked |
Exact BenchLM registry variant GPT-5.4; bulk export does not retain a complete upstream harness configuration.GPT-5.4Exact identitySource label without a registered configuration ID Canonical product: gpt-5-4 | 65.4% | Source carriersource-checkedVersion & system2026 Source-native system | BenchLM public datasets — 2026-08-01Observed Checked |
From result to context
A grounded visual reasoning benchmark focused on evidence-based question answering over real images.
The summary shows one best compatible source result per canonical product. Versions, effort settings and execution systems remain attached to the underlying records.
Scores from different versions or harnesses may not be interchangeable. The published source rows preserve those distinctions and their original units. A source result is not automatically an input to a current index.
This catalogue entry preserves available source evidence. Current indices admit only their specifically reviewed tracks and configurations.
Contamination risk: Unknown. Lifecycle: Active.