A task with its own rules.
Tool-augmented Humanity's Last Exam scores reported in DeepSeek-V4 thinking-mode evaluations.
- Organisation
- DeepSeek-AI
- Version
- 2026
Agents / Benchmark profile
Tool-augmented Humanity's Last Exam scores reported in DeepSeek-V4 thinking-mode evaluations.
Observed results
15 ranked models · higher is better · labels show rank and score
| Model | Rank | Model ID | Provider | Score | Unit |
|---|---|---|---|---|---|
| Claude Opus 5 | #1 | claude-opus-5 | Anthropic | 64.7 | percent |
| Claude Fable 5 | #2 | claude-fable-5 | Anthropic | 63 | percent |
| GLM-5.3 | #3 | glm-5-3 | Z.AI | 62.5 | percent |
| DeepSeek V4 Pro 0813 | #4 | deepseek-v4-pro-0813 | DeepSeek | 60 | percent |
| Claude Opus 4.8 | #5 | claude-opus-4-8 | Anthropic | 57.9 | percent |
| Claude Sonnet 5 | #6 | claude-sonnet-5 | Anthropic | 57.4 | percent |
| Kimi K3 | #7 | kimi-k3 | Moonshot AI | 56 | percent |
| GLM-5.2 | #8 | glm-5-2 | Z.AI | 54.7 | percent |
| Qwen3.7-Max | #9 | qwen-3-7-max | Alibaba Cloud | 53.5 | percent |
| DeepSeek V4 Flash 0731 | #10 | deepseek-v4-flash-0731 | DeepSeek | 51.5 | percent |
| DeepSeek V4 Pro | #11 | deepseek-v4-pro | DeepSeek | 48.2 | percent |
| Agents-A1 | #12 | agents-a1 | InternScience | 47.6 | percent |
| Step 3.7 Flash | #13 | step-3-7-flash | StepFun | 47.2 | percent |
| DeepSeek V4 Flash | #14 | deepseek-v4-flash | DeepSeek | 45.1 | percent |
| Nemotron 3 Ultra | #15 | nemotron-3-ultra | NVIDIA | 37.4 | percent |
One best compatible score per canonical product · higher is better
| Rank | Model | Weights | Evidence | Score |
|---|---|---|---|---|
| 01 | Claude Opus 5Anthropic | Closed weights | Public reference | 64.7% |
| 02 | Claude Fable 5Anthropic | Closed weights | Public reference | 63% |
| 03 | GLM-5.3Z.AI | Open announced | Public reference | 62.5% |
| 04 | DeepSeek V4 Pro 0813DeepSeek | Open weights | Public reference | 60% |
| 05 | Claude Opus 4.8Anthropic | Closed weights | Public reference | 57.9% |
| 06 | Claude Sonnet 5Anthropic | Closed weights | Source carrier | 57.4% |
| 07 | Kimi K3Moonshot AI | Open weights | Public reference | 56% |
| 08 | GLM-5.2Z.AI | Open weights | Public reference | 54.7% |
| 09 | Qwen3.7-MaxAlibaba Cloud | Closed weights | Source carrier | 53.5% |
| 10 | DeepSeek V4 Flash 0731DeepSeek | Open weights | Public reference | 51.5% |
Each distinct published measurement is retained. Same-snapshot canonical and configuration projections appear once.
| Published model / configuration | Score | Evidence & protocol | Source & dates |
|---|---|---|---|
GPT-5.6 Sol as published by Z.AIGPT-5.6 SolExact identitygpt-5-6-sol-max Canonical product: gpt-5-6-sol | 64.5% | Public referenceprovider-reportedVersion & systemwith tools Z.AI GLM-5.3 official comparison table | GLM-5.3: Frontier Coding with Emergent Cyber CapabilitiesObserved Checked |
Claude Fable 5 w/ fallback as published by Z.AIClaude Fable 5Exact identityclaude-fable-5-max Canonical product: claude-fable-5 | 63.9% | Public referenceprovider-reportedVersion & systemwith tools Z.AI GLM-5.3 official comparison table | GLM-5.3: Frontier Coding with Emergent Cyber CapabilitiesObserved Checked |
GLM-5.3 (max)GLM-5.3Exact identityglm-5-3-max Canonical product: glm-5-3 | 62.5% | Public referenceprovider-reportedVersion & systemwith tools Z.AI GLM-5.3 official comparison table | GLM-5.3: Frontier Coding with Emergent Cyber CapabilitiesObserved Checked |
DeepSeek V4 Pro 0813 as published by Z.AIDeepSeek V4 Pro 0813Exact identitydeepseek-v4-pro-0813-max Canonical product: deepseek-v4-pro-0813 | 60% | Public referenceprovider-reportedVersion & systemwith tools Z.AI GLM-5.3 official comparison table | GLM-5.3: Frontier Coding with Emergent Cyber CapabilitiesObserved Checked |
Kimi K3 as published by Z.AIKimi K3Exact identitykimi-k3-max Canonical product: kimi-k3 | 59.8% | Public referenceprovider-reportedVersion & systemwith tools Z.AI GLM-5.3 official comparison table | GLM-5.3: Frontier Coding with Emergent Cyber CapabilitiesObserved Checked |
Claude Opus 4.8 as published by Z.AIClaude Opus 4.8Exact identityclaude-opus-4-8-max Canonical product: claude-opus-4-8 | 57.9% | Public referenceprovider-reportedVersion & systemwith tools Z.AI GLM-5.3 official comparison table | GLM-5.3: Frontier Coding with Emergent Cyber CapabilitiesObserved Checked |
Qwen3.8 Max as published by Z.AIQwen3.8 MaxExact identityqwen-3-8-max-xhigh Canonical product: qwen-3-8-max | 56.2% | Public referenceprovider-reportedVersion & systemwith tools Z.AI GLM-5.3 official comparison table | GLM-5.3: Frontier Coding with Emergent Cyber CapabilitiesObserved Checked |
GLM-5.2 as published by Z.AIGLM-5.2Exact identityglm-5-2-max Canonical product: glm-5-2 | 54.7% | Public referenceprovider-reportedVersion & systemwith tools Z.AI GLM-5.3 official comparison table | GLM-5.3: Frontier Coding with Emergent Cyber CapabilitiesObserved Checked |
Fable 5 (with fallback)Claude Fable 5Exact identityclaude-fable-5-deepseek-0813-release-unspecified Canonical product: claude-fable-5 | 63% | Public referenceprovider-reportedVersion & systemAugust 2026, with tools DeepSeek V4 Pro 0813 official GA comparison table | DeepSeek permanent refresh sourceObserved Checked |
DeepSeek V4 Pro 0813DeepSeek V4 Pro 0813Exact identitydeepseek-v4-pro-0813-deepseek-0813-release-unspecified Canonical product: deepseek-v4-pro-0813 | 60% | Public referenceprovider-reportedVersion & systemAugust 2026, with tools DeepSeek V4 Pro 0813 official GA comparison table | DeepSeek permanent refresh sourceObserved Checked |
From result to context
Tool-augmented Humanity's Last Exam scores reported in DeepSeek-V4 thinking-mode evaluations.
The summary shows one best compatible source result per canonical product. Versions, effort settings and execution systems remain attached to the underlying records.
Scores from different versions or harnesses may not be interchangeable. The published source rows preserve those distinctions and their original units. A source result is not automatically an input to a current index.
This catalogue entry preserves available source evidence. Current indices admit only their specifically reviewed tracks and configurations.
Contamination risk: Unknown. Lifecycle: Active.