A task with its own rules.
Measures reasoning over charts in scientific papers.
- Organisation
- CharXiv
- Version
- May 2026
Multimodal / Benchmark profile
Measures reasoning over charts in scientific papers.
Observed results
20 ranked models · higher is better · labels show rank and score
| Model | Rank | Model ID | Provider | Score | Unit |
|---|---|---|---|---|---|
| Claude Mythos 5 | #1 | claude-mythos-5 | Anthropic | 93.5 | percent |
| Kimi K3 | #2 | kimi-k3 | Moonshot AI | 91.3 | percent |
| Claude Opus 4.7 | #3 | claude-opus-4-7 | Anthropic | 91 | percent |
| Qwen3.8-Flash-Next | #4 | qwen-3-8-flash-next | Alibaba Cloud | 90.6 | percent |
| Qwen3.8-27B | #5 | qwen-3-8-27b | Alibaba Cloud | 90.2 | percent |
| Claude Opus 4.8 | #6 | claude-opus-4-8 | Anthropic | 89.9 | percent |
| Gemini 3.6 Flash | #7 | gemini-3-6-flash | 89.4 | percent | |
| GLM-5.3-Flash | #8 | glm-5-3-flash | Z.AI | 89.4 | percent |
| Gemini 3.7 Flash | #9 | gemini-3-7-flash | 88.7 | percent | |
| Muse Spark 1.1 | #10 | muse-spark-1-1 | Meta | 88.4 | percent |
| Claude Sonnet 5 | #11 | claude-sonnet-5 | Anthropic | 88.3 | percent |
| Muse Spark 1.2 | #12 | muse-spark-1-2 | Meta | 87.6 | percent |
| Sakana Fugu-Ultra | #13 | sakana-fugu-ultra | Sakana AI | 86.6 | percent |
| Muse Spark | #14 | muse-spark | Meta | 86.4 | percent |
| GPT-5.6 Terra | #15 | gpt-5-6-terra | OpenAI | 85.9 | percent |
| Qwen3.7-Plus | #16 | qwen-3-7-plus | Alibaba Cloud | 85.9 | percent |
| Sakana Fugu | #17 | sakana-fugu | Sakana AI | 85.1 | percent |
| Gemini 3.5 Flash | #18 | gemini-3-5-flash | 84.2 | percent | |
| GPT-5.5 | #19 | gpt-5-5 | OpenAI | 84.1 | percent |
| Gemini 3.1 Pro Preview | #20 | gemini-3-1-pro-preview | 83.3 | percent | |
| GPT-5.4 | #21 | gpt-5-4 | OpenAI | 82.8 | percent |
| GPT-5.6 Luna | #22 | gpt-5-6-luna | OpenAI | 82.7 | percent |
| GPT-5.2 | #23 | gpt-5-2 | OpenAI | 82.1 | percent |
| Inkling | #24 | inkling | Thinking Machines Lab | 82 | percent |
| Grok 4.5 | #25 | grok-4-5 | xAI | 81.6 | percent |
| Qwen3.6 Plus | #26 | qwen3-6-plus | Alibaba Cloud | 81.5 | percent |
| Inkling-Small | #27 | inkling-small | Thinking Machines Lab | 81.3 | percent |
| MiMo-V2.5 | #28 | mimo-v2-5 | Xiaomi | 81 | percent |
| Qwen3.5 397B A17B | #29 | qwen3-5-397b | Alibaba Cloud | 80.8 | percent |
| Kimi K2.6 | #30 | kimi-k2-6 | Moonshot AI | 80.4 | percent |
| Muse Glimmer 30B | #31 | muse-glimmer-30b | Meta | 78.8 | percent |
| Qwen3.6 27B | #32 | qwen3-6-27b | Alibaba Cloud | 78.4 | percent |
| Qwen3.6-35B-A3B | #33 | qwen3-6-35b-a3b | Alibaba Cloud | 78 | percent |
| Claude Sonnet 4.6 | #34 | claude-sonnet-4-6 | Anthropic | 77.4 | percent |
| Qwen3.5-122B-A10B | #35 | qwen3-5-122b-a10b | Alibaba Cloud | 77.2 | percent |
| Nemotron 3 Nano Omni 30B A3B | #36 | nemotron-3-nano-omni-30b-a3b | NVIDIA | 76.25 | percent |
| Gemini 3.1 Flash-Lite | #37 | gemini-3-1-flash-lite | 73.2 | percent | |
| Claude Opus 4.5 | #38 | claude-opus-4-5 | Anthropic | 68.5 | percent |
| Grok 4.20 | #39 | grok-4-20 | xAI | 60.9 | percent |
| Command A+ | #40 | command-a-plus | Cohere | 52.7 | percent |
One best compatible score per canonical product · higher is better
| Rank | Model | Weights | Evidence | Score |
|---|---|---|---|---|
| 01 | Claude Mythos 5Anthropic | Closed weights | Source carrier | 93.5% |
| 02 | Kimi K3Moonshot AI | Open weights | Direct source result | 91.3% |
| 03 | Claude Opus 4.7Anthropic | Closed weights | Source carrier | 91% |
| 04 | Qwen3.8-Flash-NextAlibaba Cloud | Open weights | Direct source result | 90.6% |
| 05 | Qwen3.8-27BAlibaba Cloud | Open weights | Public reference | 90.2% |
| 06 | Claude Opus 4.8Anthropic | Closed weights | Source carrier | 89.9% |
| 07 | Gemini 3.6 FlashGoogle | Closed weights | Public reference | 89.4% |
| 08 | GLM-5.3-FlashZ.AI | Open weights | Direct source result | 89.4% |
| 09 | Gemini 3.7 FlashGoogle | Closed weights | Public reference | 88.7% |
| 10 | Muse Spark 1.1Meta | Closed weights | Direct source result | 88.4% |
Each distinct published measurement is retained. Same-snapshot canonical and configuration projections appear once.
| Published model / configuration | Score | Evidence & protocol | Source & dates |
|---|---|---|---|
Qwen3.8-Flash-Next (xhigh default thinking configuration)Qwen3.8-Flash-NextExact identityqwen-3-8-flash-next-xhigh Canonical product: qwen-3-8-flash-next | 90.6% | Direct source resultsource-checkedVersion & systemRQ With CI system:qwen3-8-flash-next:charxiv-with-ci | Qwen3.8-Flash-Next current official launch pageObserved Checked |
Qwen3.8-27B (xhigh default thinking configuration) as published in Qwen's Qwen3.8-Flash-Next comparison tableQwen3.8-27BExact identityqwen-3-8-27b-xhigh Canonical product: qwen-3-8-27b | 90.2% | Relative comparisonprovider-reportedVersion & systemRQ With CI system:qwen3-8-flash-next-comparison:cell:vision:charxiv-with-ci:1 | Qwen3.8-Flash-Next launch and official provider evaluationsObserved Checked |
GLM-5.3-Flash (max effort, documented default)GLM-5.3-FlashExact identityglm-5-3-flash-max Canonical product: glm-5-3-flash | 89.4% | Direct source resultsource-checkedVersion & systemMay 2026 system:glm-5-3-flash-launch:charxiv:0 | GLM-5.3-Flash individual evaluationsObserved Checked |
Qwen3.7-Plus (provider-published configuration) as published in Qwen's Qwen3.8-Flash-Next comparison tableQwen3.7-PlusExact identityqwen-3-7-plus-unspecified Canonical product: qwen-3-7-plus | 85.9% | Relative comparisonprovider-reportedVersion & systemRQ With CI system:qwen3-8-flash-next-comparison:cell:vision:charxiv-with-ci:2 | Qwen3.8-Flash-Next launch and official provider evaluationsObserved Checked |
Qwen3.7-Plus (provider-published configuration) as published in Qwen's Qwen3.8-Flash-Next comparison tableQwen3.7-PlusExact identityqwen-3-7-plus-unspecified Canonical product: qwen-3-7-plus | 85.8% | Relative comparisonprovider-reportedVersion & systemRQ Without CI system:qwen3-8-flash-next-comparison:cell:vision:charxiv-no-ci:2 | Qwen3.8-Flash-Next launch and official provider evaluationsObserved Checked |
Qwen3.8-27B (xhigh default thinking configuration) as published in Qwen's Qwen3.8-Flash-Next comparison tableQwen3.8-27BExact identityqwen-3-8-27b-xhigh Canonical product: qwen-3-8-27b | 83.7% | Relative comparisonprovider-reportedVersion & systemRQ Without CI system:qwen3-8-flash-next-comparison:cell:vision:charxiv-no-ci:1 | Qwen3.8-Flash-Next launch and official provider evaluationsObserved Checked |
Claude Opus 4.6 (Max) as published in Qwen's Qwen3.8-Flash-Next comparison tableClaude Opus 4.6Exact identityclaude-opus-4-6-max Canonical product: claude-opus-4-6 | 66% | Relative comparisonprovider-reportedVersion & systemRQ Without CI system:qwen3-8-flash-next-comparison:cell:vision:charxiv-no-ci:3 | Qwen3.8-Flash-Next launch and official provider evaluationsObserved Checked |
Claude Opus 5 (max) as published in Meta's comparison tableClaude Opus 5Exact identityclaude-opus-5-max Canonical product: claude-opus-5 | 89.3% | Relative comparisonprovider-reportedVersion & systemMay 2026 Meta evaluation container | Multimodal Intelligence of Muse Spark 1.2Observed Checked |
Gemini 3.7 Flash (high) as published in Meta's comparison tableGemini 3.7 FlashExact identitygemini-3-7-flash-high Canonical product: gemini-3-7-flash | 88.7% | Relative comparisonprovider-reportedVersion & systemMay 2026 Meta evaluation container | Multimodal Intelligence of Muse Spark 1.2Observed Checked |
Muse Spark 1.1 (xhigh) as published in Meta's comparison tableMuse Spark 1.1Exact identitymuse-spark-1-1-xhigh Canonical product: muse-spark-1-1 | 88.4% | Relative comparisonprovider-reportedVersion & systemMay 2026 Meta evaluation container | Multimodal Intelligence of Muse Spark 1.2Observed Checked |
From result to context
Measures reasoning over charts in scientific papers.
The summary shows one best compatible source result per canonical product. Versions, effort settings and execution systems remain attached to the underlying records.
Scores from different versions or harnesses may not be interchangeable. The published source rows preserve those distinctions and their original units. A source result is not automatically an input to a current index.
This catalogue entry preserves available source evidence. Current indices admit only their specifically reviewed tracks and configurations.
Contamination risk: Unknown. Lifecycle: Active.