A task with its own rules.
Tool-free variant of Humanity's Last Exam that isolates a model's raw frontier reasoning.
- Organisation
- OpenAI
- Version
- 2026
Knowledge / Benchmark profile
Tool-free variant of Humanity's Last Exam that isolates a model's raw frontier reasoning.
Observed results
20 ranked models · higher is better · labels show rank and score
| Model | Rank | Model ID | Provider | Score | Unit |
|---|---|---|---|---|---|
| Claude Mythos 5 | #1 | claude-mythos-5 | Anthropic | 59 | percent |
| Claude Opus 5 | #2 | claude-opus-5 | Anthropic | 56.3 | percent |
| Claude Fable 5 | #3 | claude-fable-5 | Anthropic | 53.3 | percent |
| Muse Spark 1.1 | #4 | muse-spark-1-1 | Meta | 52.2 | percent |
| Sakana Fugu-Ultra | #5 | sakana-fugu-ultra | Sakana AI | 50 | percent |
| Claude Opus 4.8 | #6 | claude-opus-4-8 | Anthropic | 49.8 | percent |
| Sakana Fugu | #7 | sakana-fugu | Sakana AI | 47.2 | percent |
| Claude Opus 4.7 | #8 | claude-opus-4-7 | Anthropic | 46.9 | percent |
| Kimi K3 | #9 | kimi-k3 | Moonshot AI | 43.5 | percent |
| Claude Sonnet 5 | #10 | claude-sonnet-5 | Anthropic | 43.2 | percent |
| GPT-5.5 | #11 | gpt-5-5 | OpenAI | 43.1 | percent |
| Muse Spark | #12 | muse-spark | Meta | 42.8 | percent |
| DeepSeek V4 Pro 0813 | #13 | deepseek-v4-pro-0813 | DeepSeek | 42.7 | percent |
| GPT-5.4 | #14 | gpt-5-4 | OpenAI | 42.7 | percent |
| GLM-5.2 | #15 | glm-5-2 | Z.AI | 40.5 | percent |
| Claude Opus 4.6 | #16 | claude-opus-4-6 | Anthropic | 40 | percent |
| DeepSeek V4 Flash 0731 | #17 | deepseek-v4-flash-0731 | DeepSeek | 37.8 | percent |
| DeepSeek V4 Pro | #18 | deepseek-v4-pro | DeepSeek | 37.7 | percent |
| DeepSeek V4 Flash | #19 | deepseek-v4-flash | DeepSeek | 34.8 | percent |
| MiMo-V2.5-Pro | #20 | mimo-v2-5-pro | Xiaomi | 34 | percent |
| Grok 4.20 | #21 | grok-4-20 | xAI | 31.6 | percent |
| Inkling-Small | #22 | inkling-small | Thinking Machines Lab | 31.6 | percent |
| Inkling | #23 | inkling | Thinking Machines Lab | 30 | percent |
| GPT-5.4 mini | #24 | gpt-5-4-mini | OpenAI | 28.2 | percent |
| Nemotron 3 Ultra | #25 | nemotron-3-ultra | NVIDIA | 26.7 | percent |
| GPT-5.4 nano | #26 | gpt-5-4-nano | OpenAI | 24.3 | percent |
| Gemma 4 31B | #27 | gemma-4-31b | 19.5 | percent | |
| NVIDIA Nemotron 3.5 Lightning 30B-A3B | #28 | nemotron-3-5-lightning-30b-a3b | NVIDIA | 11.72 | percent |
| Gemma 4 26B A4B | #29 | gemma-4-26b-a4b | 8.7 | percent | |
| Gemma 4 12B Unified | #30 | gemma-4-12b | 5.2 | percent |
One best compatible score per canonical product · higher is better
| Rank | Model | Weights | Evidence | Score |
|---|---|---|---|---|
| 01 | Claude Mythos 5Anthropic | Closed weights | Source carrier | 59% |
| 02 | Claude Opus 5Anthropic | Closed weights | Public reference | 56.3% |
| 03 | Claude Fable 5Anthropic | Closed weights | Public reference | 53.3% |
| 04 | Muse Spark 1.1Meta | Closed weights | Source carrier | 52.2% |
| 05 | Sakana Fugu-UltraSakana AI | Closed weights | Source carrier | 50% |
| 06 | Claude Opus 4.8Anthropic | Closed weights | Source carrier | 49.8% |
| 07 | Sakana FuguSakana AI | Closed weights | Source carrier | 47.2% |
| 08 | Claude Opus 4.7Anthropic | Closed weights | Source carrier | 46.9% |
| 09 | Kimi K3Moonshot AI | Open weights | Source carrier | 43.5% |
| 10 | Claude Sonnet 5Anthropic | Closed weights | Source carrier | 43.2% |
Each distinct published measurement is retained. Same-snapshot canonical and configuration projections appear once.
| Published model / configuration | Score | Evidence & protocol | Source & dates |
|---|---|---|---|
Fable 5 (with fallback)Claude Fable 5Exact identityclaude-fable-5-deepseek-0813-release-unspecified Canonical product: claude-fable-5 | 53.3% | Public referenceprovider-reportedVersion & systemAugust 2026, no tools DeepSeek V4 Pro 0813 official GA comparison table | DeepSeek permanent refresh sourceObserved Checked |
Claude Opus 4.8Claude Opus 4.8Exact identityclaude-opus-4-8-deepseek-0813-release-unspecified Canonical product: claude-opus-4-8 | 49.8% | Public referenceprovider-reportedVersion & systemAugust 2026, no tools DeepSeek V4 Pro 0813 official GA comparison table | DeepSeek permanent refresh sourceObserved Checked |
Kimi K3Kimi K3Exact identitykimi-k3-deepseek-0813-release-unspecified Canonical product: kimi-k3 | 43.5% | Public referenceprovider-reportedVersion & systemAugust 2026, no tools DeepSeek V4 Pro 0813 official GA comparison table | DeepSeek permanent refresh sourceObserved Checked |
DeepSeek V4 Pro 0813DeepSeek V4 Pro 0813Exact identitydeepseek-v4-pro-0813-deepseek-0813-release-unspecified Canonical product: deepseek-v4-pro-0813 | 42.7% | Public referenceprovider-reportedVersion & systemAugust 2026, no tools DeepSeek V4 Pro 0813 official GA comparison table | DeepSeek permanent refresh sourceObserved Checked |
GLM-5.2GLM-5.2Exact identityglm-5-2-deepseek-0813-release-unspecified Canonical product: glm-5-2 | 40.5% | Public referenceprovider-reportedVersion & systemAugust 2026, no tools DeepSeek V4 Pro 0813 official GA comparison table | DeepSeek permanent refresh sourceObserved Checked |
DeepSeek V4 Flash 0731DeepSeek V4 Flash 0731Exact identitydeepseek-v4-flash-0731-deepseek-0813-release-unspecified Canonical product: deepseek-v4-flash-0731 | 37.8% | Public referenceprovider-reportedVersion & systemAugust 2026, no tools DeepSeek V4 Pro 0813 official GA comparison table | DeepSeek permanent refresh sourceObserved Checked |
DeepSeek V4 Pro PreviewDeepSeek V4 ProExact identitydeepseek-v4-pro-deepseek-0813-release-unspecified Canonical product: deepseek-v4-pro | 37.7% | Public referenceprovider-reportedVersion & systemAugust 2026, no tools DeepSeek V4 Pro 0813 official GA comparison table | DeepSeek permanent refresh sourceObserved Checked |
DeepSeek V4 Flash PreviewDeepSeek V4 FlashExact identitydeepseek-v4-flash-deepseek-0813-release-unspecified Canonical product: deepseek-v4-flash | 34.8% | Public referenceprovider-reportedVersion & systemAugust 2026, no tools DeepSeek V4 Pro 0813 official GA comparison table | DeepSeek permanent refresh sourceObserved Checked |
NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16; NVIDIA release evaluation; temperature=1.0; top_p=0.95NVIDIA Nemotron 3.5 Lightning 30B-A3BExact identitynemotron-3-5-lightning-30b-a3b-default Canonical product: nemotron-3-5-lightning-30b-a3b | 11.7% | Public referenceprovider-reportedVersion & system2026 NVIDIA NeMo Gym / NeMo Evaluator SDK consistent release harness | NVIDIA Nemotron 3.5 Lightning 30B-A3B BF16 model cardObserved Checked |
Exact BenchLM registry variant Claude Mythos 5; bulk export does not retain a complete upstream harness configuration.Claude Mythos 5Exact identitySource label without a registered configuration ID Canonical product: claude-mythos-5 | 59% | Source carriersource-checkedVersion & system2026 Source-native system | BenchLM public datasets — 2026-08-01Observed Checked |
From result to context
Tool-free variant of Humanity's Last Exam that isolates a model's raw frontier reasoning.
The summary shows one best compatible source result per canonical product. Versions, effort settings and execution systems remain attached to the underlying records.
Scores from different versions or harnesses may not be interchangeable. The published source rows preserve those distinctions and their original units. A source result is not automatically an input to a current index.
This catalogue entry preserves available source evidence. Current indices admit only their specifically reviewed tracks and configurations.
Contamination risk: Unknown. Lifecycle: Active.