A task with its own rules.
A harder multimodal benchmark for frontier models that combines text with images, diagrams, charts, and academic visual reasoning tasks.
- Organisation
- MMMU-Pro authors
- Version
- 2024
Multimodal / Benchmark profile
A harder multimodal benchmark for frontier models that combines text with images, diagrams, charts, and academic visual reasoning tasks.
Observed results
20 ranked models · higher is better · labels show rank and score
| Model | Rank | Model ID | Provider | Score | Unit |
|---|---|---|---|---|---|
| GPT-5.4 | #1 | gpt-5-4 | OpenAI | 94 | percent |
| Gemini 3.5 Flash | #2 | gemini-3-5-flash | 83.6 | percent | |
| GPT-5.6 Sol | #3 | gpt-5-6-sol | OpenAI | 83 | percent |
| Kimi K3 | #4 | kimi-k3 | Moonshot AI | 81.6 | percent |
| GPT-5.5 | #5 | gpt-5-5 | OpenAI | 81.2 | percent |
| GPT-5.6 Terra | #6 | gpt-5-6-terra | OpenAI | 80.7 | percent |
| Muse Spark | #7 | muse-spark | Meta | 80.4 | percent |
| GPT-5.2 | #8 | gpt-5-2 | OpenAI | 79.5 | percent |
| Kimi K2.6 | #9 | kimi-k2-6 | Moonshot AI | 79.4 | percent |
| Qwen3.5 397B A17B | #10 | qwen3-5-397b | Alibaba Cloud | 79 | percent |
| Qwen3.7-Plus | #11 | qwen-3-7-plus | Alibaba Cloud | 79 | percent |
| Qwen3.6 Plus | #12 | qwen3-6-plus | Alibaba Cloud | 78.8 | percent |
| Kimi K2.5 | #13 | kimi-k2-5 | Moonshot AI | 78.5 | percent |
| GPT-5.6 Luna | #14 | gpt-5-6-luna | OpenAI | 78.4 | percent |
| Grok 4.3 | #15 | grok-4-3 | xAI | 78.1 | percent |
| MiniMax M3 | #16 | minimax-m3 | MiniMax | 78.1 | percent |
| MiMo-V2.5 | #17 | mimo-v2-5 | Xiaomi | 77.9 | percent |
| Claude Opus 4.6 | #18 | claude-opus-4-6 | Anthropic | 77.3 | percent |
| Gemma 4 31B | #19 | gemma-4-31b | 76.9 | percent | |
| GPT-5.4 mini | #20 | gpt-5-4-mini | OpenAI | 76.6 | percent |
| Qwen3.6 27B | #21 | qwen3-6-27b | Alibaba Cloud | 75.8 | percent |
| Qwen3.6-35B-A3B | #22 | qwen3-6-35b-a3b | Alibaba Cloud | 75.3 | percent |
| Grok 4.20 | #23 | grok-4-20 | xAI | 75.2 | percent |
| Inkling-Small | #24 | inkling-small | Thinking Machines Lab | 74 | percent |
| Gemma 4 26B A4B | #25 | gemma-4-26b-a4b | 73.8 | percent | |
| Inkling | #26 | inkling | Thinking Machines Lab | 73.5 | percent |
| Interfaze Beta | #27 | interfaze-beta | Interfaze | 71.1 | percent |
| Claude Opus 4.5 | #28 | claude-opus-4-5 | Anthropic | 70.6 | percent |
| Gemma 4 12B Unified | #29 | gemma-4-12b | 69.1 | percent | |
| GPT-5.4 nano | #30 | gpt-5-4-nano | OpenAI | 66.1 | percent |
| Command A+ | #31 | command-a-plus | Cohere | 63 | percent |
One best compatible score per canonical product · higher is better
| Rank | Model | Weights | Evidence | Score |
|---|---|---|---|---|
| 01 | GPT-5.4OpenAI | Closed weights | Source carrier | 94% |
| 02 | Gemini 3.5 FlashGoogle | Closed weights | Source carrier | 83.6% |
| 03 | GPT-5.6 SolOpenAI | Closed weights | Source carrier | 83% |
| 04 | Kimi K3Moonshot AI | Open weights | Source carrier | 81.6% |
| 05 | GPT-5.5OpenAI | Closed weights | Source carrier | 81.2% |
| 06 | GPT-5.6 TerraOpenAI | Closed weights | Source carrier | 80.7% |
| 07 | Muse SparkMeta | Closed weights | Source carrier | 80.4% |
| 08 | GPT-5.2OpenAI | Closed weights | Source carrier | 79.5% |
| 09 | Kimi K2.6Moonshot AI | Open weights | Source carrier | 79.4% |
| 10 | Qwen3.5 397B A17BAlibaba Cloud | Open weights | Source carrier | 79% |
Each distinct published measurement is retained. Same-snapshot canonical and configuration projections appear once.
| Published model / configuration | Score | Evidence & protocol | Source & dates |
|---|---|---|---|
Exact BenchLM registry variant GPT-5.4 Pro; bulk export does not retain a complete upstream harness configuration.GPT-5.4Exact identitygpt-5-4-pro-default-medium Canonical product: gpt-5-4 | 94% | Source carriersource-checkedVersion & system2024 Source-native system | BenchLM public datasets — 2026-08-01Observed Checked |
Exact BenchLM registry variant Gemini 3.5 Flash; bulk export does not retain a complete upstream harness configuration.Gemini 3.5 FlashExact identitySource label without a registered configuration ID Canonical product: gemini-3-5-flash | 83.6% | Source carriersource-checkedVersion & system2024 Source-native system | BenchLM public datasets — 2026-08-01Observed Checked |
Exact BenchLM registry variant GPT-5.6 Sol; bulk export does not retain a complete upstream harness configuration.GPT-5.6 SolExact identitySource label without a registered configuration ID Canonical product: gpt-5-6-sol | 83% | Source carriersource-checkedVersion & system2024 Source-native system | BenchLM public datasets — 2026-08-01Observed Checked |
Exact BenchLM registry variant Kimi K3; bulk export does not retain a complete upstream harness configuration.Kimi K3Exact identitySource label without a registered configuration ID Canonical product: kimi-k3 | 81.6% | Source carriersource-checkedVersion & system2024 Source-native system | BenchLM public datasets — 2026-08-01Observed Checked |
Exact BenchLM registry variant GPT-5.4; bulk export does not retain a complete upstream harness configuration.GPT-5.4Exact identitySource label without a registered configuration ID Canonical product: gpt-5-4 | 81.2% | Source carriersource-checkedVersion & system2024 Source-native system | BenchLM public datasets — 2026-08-01Observed Checked |
Exact BenchLM registry variant GPT-5.5; bulk export does not retain a complete upstream harness configuration.GPT-5.5Exact identitySource label without a registered configuration ID Canonical product: gpt-5-5 | 81.2% | Source carriersource-checkedVersion & system2024 Source-native system | BenchLM public datasets — 2026-08-01Observed Checked |
Exact BenchLM registry variant GPT-5.6 Terra; bulk export does not retain a complete upstream harness configuration.GPT-5.6 TerraExact identitySource label without a registered configuration ID Canonical product: gpt-5-6-terra | 80.7% | Source carriersource-checkedVersion & system2024 Source-native system | BenchLM public datasets — 2026-08-01Observed Checked |
Exact BenchLM registry variant Muse Spark; bulk export does not retain a complete upstream harness configuration.Muse SparkExact identitySource label without a registered configuration ID Canonical product: muse-spark | 80.4% | Source carriersource-checkedVersion & system2024 Source-native system | BenchLM public datasets — 2026-08-01Observed Checked |
Exact BenchLM registry variant GPT-5.2; bulk export does not retain a complete upstream harness configuration.GPT-5.2Exact identitySource label without a registered configuration ID Canonical product: gpt-5-2 | 79.5% | Source carriersource-checkedVersion & system2024 Source-native system | BenchLM public datasets — 2026-08-01Observed Checked |
Exact BenchLM registry variant Kimi K2.6; bulk export does not retain a complete upstream harness configuration.Kimi K2.6Exact identitySource label without a registered configuration ID Canonical product: kimi-k2-6 | 79.4% | Source carriersource-checkedVersion & system2024 Source-native system | BenchLM public datasets — 2026-08-01Observed Checked |
From result to context
A harder multimodal benchmark for frontier models that combines text with images, diagrams, charts, and academic visual reasoning tasks.
The summary shows one best compatible source result per canonical product. Versions, effort settings and execution systems remain attached to the underlying records.
Scores from different versions or harnesses may not be interchangeable. The published source rows preserve those distinctions and their original units. A source result is not automatically an input to a current index.
This catalogue entry preserves available source evidence. Current indices admit only their specifically reviewed tracks and configurations.
Contamination risk: Unknown. Lifecycle: Active.