A task with its own rules.
A more challenging multiple-choice knowledge and reasoning benchmark derived from MMLU.
- Organisation
- TIGER-Lab
- Version
- Current / rolling
Research / Benchmark profile
A more challenging multiple-choice knowledge and reasoning benchmark derived from MMLU.
Observed results
20 ranked models · higher is better · labels show rank and score
| Model | Rank | Model ID | Provider | Score | Unit |
|---|---|---|---|---|---|
| Qwen3.7-Max | #1 | qwen-3-7-max | Alibaba Cloud | 89.6 | percent |
| Claude Opus 4.5 | #2 | claude-opus-4-5 | Anthropic | 89.5 | percent |
| Qwen3.6 Plus | #3 | qwen3-6-plus | Alibaba Cloud | 88.5 | percent |
| Qwen3.7-Plus | #4 | qwen-3-7-plus | Alibaba Cloud | 88.5 | percent |
| Qwen3.5 397B A17B | #5 | qwen3-5-397b | Alibaba Cloud | 87.8 | percent |
| Kimi K2.5 | #6 | kimi-k2-5 | Moonshot AI | 87.1 | percent |
| Nemotron 3 Ultra | #7 | nemotron-3-ultra | NVIDIA | 86.8 | percent |
| Qwen3.5 122B | #8 | qwen3-5-122b | Alibaba Cloud | 86.7 | percent |
| Qwen3.6 27B | #9 | qwen3-6-27b | Alibaba Cloud | 86.2 | percent |
| Solar Open 2 250B | #10 | solar-open2-250b | Upstage | 86.2 | percent |
| Qwen3.5 27B | #11 | qwen3-5-27b | Alibaba Cloud | 86.1 | percent |
| GLM-5 | #12 | glm-5 | Z.AI | 85.7 | percent |
| Qwen3.5 35B | #13 | qwen3-5-35b | Alibaba Cloud | 85.3 | percent |
| Gemma 4 31B | #14 | gemma-4-31b | 85.2 | percent | |
| GLM-4.7 | #15 | glm-4-7 | Z.AI | 84.3 | percent |
| DeepSeek V4 Flash | #16 | deepseek-v4-flash | DeepSeek | 83 | percent |
| DeepSeek V4 Pro | #17 | deepseek-v4-pro | DeepSeek | 82.9 | percent |
| Gemma 4 26B | #18 | gemma-4-26b | 82.6 | percent | |
| Claude Opus 4.6 | #19 | claude-opus-4-6 | Anthropic | 82 | percent |
| NVIDIA Nemotron 3.5 Lightning 30B-A3B | #20 | nemotron-3-5-lightning-30b-a3b | NVIDIA | 81.94 | percent |
| Claude Sonnet 4.6 | #21 | claude-sonnet-4-6 | Anthropic | 79.2 | percent |
| Gemma 4 12B Unified | #22 | gemma-4-12b | 77.2 | percent |
One best compatible score per canonical product · higher is better
| Rank | Model | Weights | Evidence | Score |
|---|---|---|---|---|
| 01 | Qwen3.7-MaxAlibaba Cloud | Closed weights | Source carrier | 89.6% |
| 02 | Claude Opus 4.5Anthropic | Closed weights | Source carrier | 89.5% |
| 03 | Qwen3.6 PlusAlibaba Cloud | Closed weights | Source carrier | 88.5% |
| 04 | Qwen3.7-PlusAlibaba Cloud | Closed weights | Source carrier | 88.5% |
| 05 | Qwen3.5 397B A17BAlibaba Cloud | Open weights | Direct source result | 87.8% |
| 06 | Kimi K2.5Moonshot AI | Open weights | Source carrier | 87.1% |
| 07 | Nemotron 3 UltraNVIDIA | Open weights | Direct source result | 86.8% |
| 08 | Qwen3.5 122BAlibaba Cloud | Open weights | Direct source result | 86.7% |
| 09 | Qwen3.6 27BAlibaba Cloud | Open weights | Direct source result | 86.2% |
| 10 | Solar Open 2 250BUpstage | Open weights | Direct source result | 86.2% |
Each distinct published measurement is retained. Same-snapshot canonical and configuration projections appear once.
| Published model / configuration | Score | Evidence & protocol | Source & dates |
|---|---|---|---|
Solar Open 2 250B (high)Solar Open 2 250BExact identitysolar-open2-250b-high Canonical product: solar-open2-250b | 86.2% | Direct source resultprovider-reportedVersion & systemstandard Solar Open 2 card English table | Solar Open 2 250B model cardObserved Checked |
DeepSeek V4 Flash max as published by UpstageDeepSeek V4 FlashExact identitydeepseek-v4-flash-max Canonical product: deepseek-v4-flash | 85.9% | Relative comparisonprovider-reportedVersion & systemstandard Solar Open 2 card English table | Solar Open 2 250B model cardObserved Checked |
MiMo-V2.5 as published by UpstageMiMo-V2.5Exact identitymimo-v2-5-solar-open2-unspecified Canonical product: mimo-v2-5 | 84.6% | Relative comparisonprovider-reportedVersion & systemstandard Solar Open 2 card English table | Solar Open 2 250B model cardObserved Checked |
Mistral Medium 3.5 as published by UpstageMistral Medium 3.5Exact identitymistral-medium-3-5-high Canonical product: mistral-medium-3-5 | 81.2% | Relative comparisonprovider-reportedVersion & systemstandard Solar Open 2 card English table | Solar Open 2 250B model cardObserved Checked |
Solar Open 100B as published by UpstageSolar Open 100B (Reasoning)Exact identitysolar-open-100b-reasoning-high Canonical product: solar-open-100b-reasoning | 80.4% | Relative comparisonprovider-reportedVersion & systemstandard Solar Open 2 card English table | Solar Open 2 250B model cardObserved Checked |
Command A+ as published by UpstageCommand A+Exact identitycommand-a-plus-solar-open2-unspecified Canonical product: command-a-plus | 79% | Relative comparisonprovider-reportedVersion & systemstandard Solar Open 2 card English table | Solar Open 2 250B model cardObserved Checked |
NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16; NVIDIA release evaluation; temperature=1.0; top_p=0.95NVIDIA Nemotron 3.5 Lightning 30B-A3BExact identitynemotron-3-5-lightning-30b-a3b-default Canonical product: nemotron-3-5-lightning-30b-a3b | 81.9% | Direct source resultprovider-reportedVersion & systemCurrent NVIDIA NeMo Gym / NeMo Evaluator SDK consistent release harness | NVIDIA Nemotron 3.5 Lightning 30B-A3B BF16 model cardObserved Checked |
Exact model variant as listed on the BenchLM public mmlu-pro leaderboard page (verified 2026-07-21).Qwen3.5 397B A17BExact identitySource label without a registered configuration ID Canonical product: qwen3-5-397b | 87.8% | Direct source resultsource-checkedVersion & systemCurrent BenchLM aggregated public evaluation | MMLU-Pro Leaderboard & Scores — July 2026Observed Checked |
Exact model variant as listed on the BenchLM public mmlu-pro leaderboard page (verified 2026-07-21).Nemotron 3 UltraExact identitySource label without a registered configuration ID Canonical product: nemotron-3-ultra | 86.8% | Direct source resultsource-checkedVersion & systemCurrent BenchLM aggregated public evaluation | MMLU-Pro Leaderboard & Scores — July 2026Observed Checked |
Exact model variant as listed on the BenchLM public mmlu-pro leaderboard page (verified 2026-07-21).Qwen3.5 122BExact identitySource label without a registered configuration ID Canonical product: qwen3-5-122b | 86.7% | Direct source resultsource-checkedVersion & systemCurrent BenchLM aggregated public evaluation | MMLU-Pro Leaderboard & Scores — July 2026Observed Checked |
From result to context
A more challenging multiple-choice knowledge and reasoning benchmark derived from MMLU.
The summary shows one best compatible source result per canonical product. Versions, effort settings and execution systems remain attached to the underlying records.
Scores from different versions or harnesses may not be interchangeable. The published source rows preserve those distinctions and their original units. A source result is not automatically an input to a current index.
This catalogue entry preserves available source evidence. Current indices admit only their specifically reviewed tracks and configurations.
Contamination risk: Unknown. Lifecycle: Active.