A task with its own rules.
A challenging mathematical reasoning benchmark reported in DeepSeek-V4 model evaluations.
- Organisation
- DeepSeek-AI
- Version
- 2026
Mathematics / Benchmark profile
A challenging mathematical reasoning benchmark reported in DeepSeek-V4 model evaluations.
Observed results
6 ranked models · higher is better · labels show rank and score
| Model | Rank | Model ID | Provider | Score | Unit |
|---|---|---|---|---|---|
| Qwen3.7-Max | #1 | qwen-3-7-max | Alibaba Cloud | 90 | percent |
| DeepSeek V4 Pro | #2 | deepseek-v4-pro | DeepSeek | 89.8 | percent |
| DeepSeek V4 Flash | #3 | deepseek-v4-flash | DeepSeek | 88.4 | percent |
| Qwen3.7-Plus | #4 | qwen-3-7-plus | Alibaba Cloud | 86 | percent |
| LongCat-2.0 | #5 | longcat-2-0 | Meituan | 81.8 | percent |
| ZAYA1-8B | #6 | zaya1-8b | Zyphra | 59.3 | percent |
One best compatible score per canonical product · higher is better
| Rank | Model | Weights | Evidence | Score |
|---|---|---|---|---|
| 01 | Qwen3.7-MaxAlibaba Cloud | Closed weights | Source carrier | 90% |
| 02 | DeepSeek V4 ProDeepSeek | Open weights | Source carrier | 89.8% |
| 03 | DeepSeek V4 FlashDeepSeek | Open weights | Source carrier | 88.4% |
| 04 | Qwen3.7-PlusAlibaba Cloud | Closed weights | Source carrier | 86% |
| 05 | LongCat-2.0Meituan | Open weights | Public reference | 81.8% |
| 06 | ZAYA1-8BZyphra | Open weights | Source carrier | 59.3% |
Each distinct published measurement is retained. Same-snapshot canonical and configuration projections appear once.
| Published model / configuration | Score | Evidence & protocol | Source & dates |
|---|---|---|---|
Exact BenchLM registry variant Qwen3.7 Max; bulk export does not retain a complete upstream harness configuration.Qwen3.7-MaxExact identitySource label without a registered configuration ID Canonical product: qwen-3-7-max | 90% | Source carriersource-checkedVersion & system2026 Source-native system | BenchLM public datasets — 2026-08-01Observed Checked |
Exact BenchLM registry variant DeepSeek V4 Pro (Max); bulk export does not retain a complete upstream harness configuration.DeepSeek V4 ProExact identitydeepseek-v4-pro-max Canonical product: deepseek-v4-pro | 89.8% | Source carriersource-checkedVersion & system2026 Source-native system | BenchLM public datasets — 2026-08-01Observed Checked |
Exact BenchLM registry variant DeepSeek V4 Flash (Max); bulk export does not retain a complete upstream harness configuration.DeepSeek V4 FlashExact identitydeepseek-v4-flash-max Canonical product: deepseek-v4-flash | 88.4% | Source carriersource-checkedVersion & system2026 Source-native system | BenchLM public datasets — 2026-08-01Observed Checked |
Exact BenchLM registry variant DeepSeek V4 Pro (High); bulk export does not retain a complete upstream harness configuration.DeepSeek V4 ProExact identitydeepseek-v4-pro-high Canonical product: deepseek-v4-pro | 88% | Source carriersource-checkedVersion & system2026 Source-native system | BenchLM public datasets — 2026-08-01Observed Checked |
Exact BenchLM registry variant Qwen3.7 Plus; bulk export does not retain a complete upstream harness configuration.Qwen3.7-PlusExact identitySource label without a registered configuration ID Canonical product: qwen-3-7-plus | 86% | Source carriersource-checkedVersion & system2026 Source-native system | BenchLM public datasets — 2026-08-01Observed Checked |
Exact BenchLM registry variant DeepSeek V4 Flash (High); bulk export does not retain a complete upstream harness configuration.DeepSeek V4 FlashExact identitydeepseek-v4-flash-high Canonical product: deepseek-v4-flash | 85.1% | Source carriersource-checkedVersion & system2026 Source-native system | BenchLM public datasets — 2026-08-01Observed Checked |
Exact BenchLM registry variant ZAYA1-8B; bulk export does not retain a complete upstream harness configuration.ZAYA1-8BExact identitySource label without a registered configuration ID Canonical product: zaya1-8b | 59.3% | Source carriersource-checkedVersion & system2026 Source-native system | BenchLM public datasets — 2026-08-01Observed Checked |
Exact BenchLM registry variant DeepSeek V4 Flash; bulk export does not retain a complete upstream harness configuration.DeepSeek V4 FlashExact identitySource label without a registered configuration ID Canonical product: deepseek-v4-flash | 41.9% | Source carriersource-checkedVersion & system2026 Source-native system | BenchLM public datasets — 2026-08-01Observed Checked |
Exact BenchLM registry variant DeepSeek V4 Pro; bulk export does not retain a complete upstream harness configuration.DeepSeek V4 ProExact identitySource label without a registered configuration ID Canonical product: deepseek-v4-pro | 35.3% | Source carriersource-checkedVersion & system2026 Source-native system | BenchLM public datasets — 2026-08-01Observed Checked |
Exact BenchLM registry variant Qwen3.7 Max; bulk export does not retain a complete upstream harness configuration.Qwen3.7-MaxExact identitySource label without a registered configuration ID Canonical product: qwen-3-7-max | 90% | Source carriersource-checkedVersion & system2026 Source-native system | BenchLM public datasets — 2026-07-27Observed Checked |
From result to context
A challenging mathematical reasoning benchmark reported in DeepSeek-V4 model evaluations.
The summary shows one best compatible source result per canonical product. Versions, effort settings and execution systems remain attached to the underlying records.
Scores from different versions or harnesses may not be interchangeable. The published source rows preserve those distinctions and their original units. A source result is not automatically an input to a current index.
This catalogue entry preserves available source evidence. Current indices admit only their specifically reviewed tracks and configurations.
Contamination risk: Unknown. Lifecycle: Active.