A task with its own rules.
A benchmark of difficult, expert-authored mathematics problems designed for frontier systems.
- Organisation
- Epoch AI
- Version
- Current / rolling
Mathematics / Benchmark profile
A benchmark of difficult, expert-authored mathematics problems designed for frontier systems.
Observed results
20 ranked models · higher is better · labels show rank and score
| Model | Rank | Model ID | Provider | Score | Unit |
|---|---|---|---|---|---|
| GPT-5.6 Sol | #1 | gpt-5-6-sol | OpenAI | 89 | percent |
| GPT-5.6 Terra | #2 | gpt-5-6-terra | OpenAI | 84.9 | percent |
| GPT-5.6 Luna | #3 | gpt-5-6-luna | OpenAI | 78.6 | percent |
| GPT-5.5 | #4 | gpt-5-5 | OpenAI | 52.4 | percent |
| GPT-5.4 | #5 | gpt-5-4 | OpenAI | 50 | percent |
| Claude Opus 4.8 | #6 | claude-opus-4-8 | Anthropic | 47.241 | percent |
| Claude Opus 4.7 | #7 | claude-opus-4-7 | Anthropic | 43.8 | percent |
| Claude Opus 4.6 | #8 | claude-opus-4-6 | Anthropic | 40.7 | percent |
| Muse Spark | #9 | muse-spark | Meta | 39 | percent |
| Gemini 3.5 Flash | #10 | gemini-3-5-flash | 38.966 | percent | |
| Gemini 3.1 Pro Preview | #11 | gemini-3-1-pro-preview | 36.9 | percent | |
| Gemini 3 Flash | #12 | gemini-3-flash | 35.64 | percent | |
| GLM-5.1 | #13 | glm-5-1 | Z.AI | 33.448 | percent |
| Claude Sonnet 4.6 | #14 | claude-sonnet-4-6 | Anthropic | 32.4 | percent |
| GPT-5.1 | #15 | gpt-5-1 | OpenAI | 31.034 | percent |
| Kimi K2.5 | #16 | kimi-k2-5 | Moonshot AI | 27.9 | percent |
| Qwen3.6 Plus | #17 | qwen3-6-plus | Alibaba Cloud | 26.207 | percent |
| GPT-5.4 nano | #18 | gpt-5-4-nano | OpenAI | 25.86 | percent |
| Claude Opus 4.5 | #19 | claude-opus-4-5 | Anthropic | 20.69 | percent |
| GLM-5 | #20 | glm-5 | Z.AI | 16.434 | percent |
| Claude Haiku 4.5 | #21 | claude-haiku-4-5 | Anthropic | 5.903 | percent |
One best compatible score per canonical product · higher is better
| Rank | Model | Weights | Evidence | Score |
|---|---|---|---|---|
| 01 | GPT-5.6 SolOpenAI | Closed weights | Source carrier | 89% |
| 02 | GPT-5.6 TerraOpenAI | Closed weights | Source carrier | 84.9% |
| 03 | GPT-5.6 LunaOpenAI | Closed weights | Source carrier | 78.6% |
| 04 | GPT-5.5OpenAI | Closed weights | Source carrier | 52.4% |
| 05 | GPT-5.4OpenAI | Closed weights | Source carrier | 50% |
| 06 | Claude Opus 4.8Anthropic | Closed weights | Source carrier | 47.2% |
| 07 | Claude Opus 4.7Anthropic | Closed weights | Source carrier | 43.8% |
| 08 | Claude Opus 4.6Anthropic | Closed weights | Source carrier | 40.7% |
| 09 | Muse SparkMeta | Closed weights | Source carrier | 39% |
| 10 | Gemini 3.5 FlashGoogle | Closed weights | Source carrier | 39.0% |
Each distinct published measurement is retained. Same-snapshot canonical and configuration projections appear once.
| Published model / configuration | Score | Evidence & protocol | Source & dates |
|---|---|---|---|
Exact BenchLM registry variant GPT-5.6 Sol; bulk export does not retain a complete upstream harness configuration.GPT-5.6 SolExact identitySource label without a registered configuration ID Canonical product: gpt-5-6-sol | 89% | Source carriersource-checkedVersion & system2024 Source-native system | BenchLM public datasets — 2026-08-01Observed Checked |
Exact BenchLM registry variant GPT-5.6 Terra; bulk export does not retain a complete upstream harness configuration.GPT-5.6 TerraExact identitySource label without a registered configuration ID Canonical product: gpt-5-6-terra | 84.9% | Source carriersource-checkedVersion & system2024 Source-native system | BenchLM public datasets — 2026-08-01Observed Checked |
Exact BenchLM registry variant GPT-5.6 Luna; bulk export does not retain a complete upstream harness configuration.GPT-5.6 LunaExact identitySource label without a registered configuration ID Canonical product: gpt-5-6-luna | 78.6% | Source carriersource-checkedVersion & system2024 Source-native system | BenchLM public datasets — 2026-08-01Observed Checked |
Exact BenchLM registry variant GPT-5.5 Pro; bulk export does not retain a complete upstream harness configuration.GPT-5.5Exact identitygpt-5-5-pro-default-high Canonical product: gpt-5-5 | 52.4% | Source carriersource-checkedVersion & system2024 Source-native system | BenchLM public datasets — 2026-08-01Observed Checked |
Exact BenchLM registry variant GPT-5.5; bulk export does not retain a complete upstream harness configuration.GPT-5.5Exact identitySource label without a registered configuration ID Canonical product: gpt-5-5 | 51.7% | Source carriersource-checkedVersion & system2024 Source-native system | BenchLM public datasets — 2026-08-01Observed Checked |
Exact BenchLM registry variant GPT-5.4 Pro; bulk export does not retain a complete upstream harness configuration.GPT-5.4Exact identitygpt-5-4-pro-default-medium Canonical product: gpt-5-4 | 50% | Source carriersource-checkedVersion & system2024 Source-native system | BenchLM public datasets — 2026-08-01Observed Checked |
Exact BenchLM registry variant Claude Opus 4.7 (Adaptive); bulk export does not retain a complete upstream harness configuration.Claude Opus 4.7Exact identityclaude-opus-4-7-max Canonical product: claude-opus-4-7 | 43.8% | Source carriersource-checkedVersion & system2024 Source-native system | BenchLM public datasets — 2026-08-01Observed Checked |
Exact BenchLM registry variant GPT-5.6 Sol; bulk export does not retain a complete upstream harness configuration.GPT-5.6 SolExact identitySource label without a registered configuration ID Canonical product: gpt-5-6-sol | 89% | Source carriersource-checkedVersion & system2024 Source-native system | BenchLM public datasets — 2026-07-27Observed Checked |
Exact BenchLM registry variant GPT-5.6 Terra; bulk export does not retain a complete upstream harness configuration.GPT-5.6 TerraExact identitySource label without a registered configuration ID Canonical product: gpt-5-6-terra | 84.9% | Source carriersource-checkedVersion & system2024 Source-native system | BenchLM public datasets — 2026-07-27Observed Checked |
Exact BenchLM registry variant GPT-5.6 Luna; bulk export does not retain a complete upstream harness configuration.GPT-5.6 LunaExact identitySource label without a registered configuration ID Canonical product: gpt-5-6-luna | 78.6% | Source carriersource-checkedVersion & system2024 Source-native system | BenchLM public datasets — 2026-07-27Observed Checked |
From result to context
A benchmark of difficult, expert-authored mathematics problems designed for frontier systems.
The summary shows one best compatible source result per canonical product. Versions, effort settings and execution systems remain attached to the underlying records.
Scores from different versions or harnesses may not be interchangeable. The published source rows preserve those distinctions and their original units. A source result is not automatically an input to a current index.
A reviewed track from this benchmark is represented in the following current indices. Each index applies its own protocol and eligibility rules.
Contamination risk: Low. Lifecycle: Rolling.