A task with its own rules.
Harvard-MIT Mathematics Tournament February 2026 problems used as a contest math evaluation.
- Organisation
- HMMT
- Version
- Current / rolling
Mathematics / Benchmark profile
Harvard-MIT Mathematics Tournament February 2026 problems used as a contest math evaluation.
Observed results
11 ranked models · higher is better · labels show rank and score
| Model | Rank | Model ID | Provider | Score | Unit |
|---|---|---|---|---|---|
| Qwen3.7-Max | #1 | qwen-3-7-max | Alibaba Cloud | 97.1 | percent |
| Solar Open 2 250B | #2 | solar-open2-250b | Upstage | 93.9 | percent |
| Qwen3.7-Plus | #3 | qwen-3-7-plus | Alibaba Cloud | 92.9 | percent |
| GLM-5.2 | #4 | glm-5-2 | Z.AI | 92.5 | percent |
| Qwen3.6 Plus | #5 | qwen3-6-plus | Alibaba Cloud | 87.8 | percent |
| Kimi K2.5 | #6 | kimi-k2-5 | Moonshot AI | 87.1 | percent |
| GLM-5 | #7 | glm-5 | Z.AI | 86.4 | percent |
| Claude Opus 4.5 | #8 | claude-opus-4-5 | Anthropic | 85.3 | percent |
| GLM-5.1 | #9 | glm-5-1 | Z.AI | 82.6 | percent |
| DeepSeek V4 Flash | #10 | deepseek-v4-flash | DeepSeek | 40.8 | percent |
| DeepSeek V4 Pro | #11 | deepseek-v4-pro | DeepSeek | 31.7 | percent |
One best compatible score per canonical product · higher is better
| Rank | Model | Weights | Evidence | Score |
|---|---|---|---|---|
| 01 | Qwen3.7-MaxAlibaba Cloud | Closed weights | Source carrier | 97.1% |
| 02 | Solar Open 2 250BUpstage | Open weights | Direct source result | 93.9% |
| 03 | Qwen3.7-PlusAlibaba Cloud | Closed weights | Source carrier | 92.9% |
| 04 | GLM-5.2Z.AI | Open weights | Source carrier | 92.5% |
| 05 | Qwen3.6 PlusAlibaba Cloud | Closed weights | Source carrier | 87.8% |
| 06 | Kimi K2.5Moonshot AI | Open weights | Source carrier | 87.1% |
| 07 | GLM-5Z.AI | Open weights | Source carrier | 86.4% |
| 08 | Claude Opus 4.5Anthropic | Closed weights | Source carrier | 85.3% |
| 09 | GLM-5.1Z.AI | Open weights | Source carrier | 82.6% |
| 10 | DeepSeek V4 FlashDeepSeek | Open weights | Source carrier | 40.8% |
Each distinct published measurement is retained. Same-snapshot canonical and configuration projections appear once.
| Published model / configuration | Score | Evidence & protocol | Source & dates |
|---|---|---|---|
DeepSeek V4 Flash max as published by UpstageDeepSeek V4 FlashExact identitydeepseek-v4-flash-max Canonical product: deepseek-v4-flash | 94.7% | Relative comparisonprovider-reportedVersion & system2602 Solar Open 2 card English table | Solar Open 2 250B model cardObserved Checked |
Solar Open 2 250B (high)Solar Open 2 250BExact identitysolar-open2-250b-high Canonical product: solar-open2-250b | 93.9% | Direct source resultprovider-reportedVersion & system2602 Solar Open 2 card English table | Solar Open 2 250B model cardObserved Checked |
Command A+ as published by UpstageCommand A+Exact identitycommand-a-plus-solar-open2-unspecified Canonical product: command-a-plus | 73.5% | Relative comparisonprovider-reportedVersion & system2602 Solar Open 2 card English table | Solar Open 2 250B model cardObserved Checked |
Solar Open 100B as published by UpstageSolar Open 100B (Reasoning)Exact identitysolar-open-100b-reasoning-high Canonical product: solar-open-100b-reasoning | 68.9% | Relative comparisonprovider-reportedVersion & system2602 Solar Open 2 card English table | Solar Open 2 250B model cardObserved Checked |
Mistral Medium 3.5 as published by UpstageMistral Medium 3.5Exact identitymistral-medium-3-5-high Canonical product: mistral-medium-3-5 | 62.9% | Relative comparisonprovider-reportedVersion & system2602 Solar Open 2 card English table | Solar Open 2 250B model cardObserved Checked |
MiMo-V2.5 as published by UpstageMiMo-V2.5Exact identitymimo-v2-5-solar-open2-unspecified Canonical product: mimo-v2-5 | 61.4% | Relative comparisonprovider-reportedVersion & system2602 Solar Open 2 card English table | Solar Open 2 250B model cardObserved Checked |
Exact model variant as listed on the BenchLM public benchmark leaderboard page.Qwen3.7-MaxExact identitySource label without a registered configuration ID Canonical product: qwen-3-7-max | 97.1% | Source carriersource-checkedVersion & systemCurrent BenchLM aggregated public evaluation | BenchLM public benchmark leaderboardsObserved Checked |
Exact model variant as listed on the BenchLM public benchmark leaderboard page.Qwen3.7-PlusExact identitySource label without a registered configuration ID Canonical product: qwen-3-7-plus | 92.9% | Source carriersource-checkedVersion & systemCurrent BenchLM aggregated public evaluation | BenchLM public benchmark leaderboardsObserved Checked |
Exact model variant as listed on the BenchLM public benchmark leaderboard page.GLM-5.2Exact identitySource label without a registered configuration ID Canonical product: glm-5-2 | 92.5% | Source carriersource-checkedVersion & systemCurrent BenchLM aggregated public evaluation | BenchLM public benchmark leaderboardsObserved Checked |
Exact model variant as listed on the BenchLM public benchmark leaderboard page.Qwen3.6 PlusExact identitySource label without a registered configuration ID Canonical product: qwen3-6-plus | 87.8% | Source carriersource-checkedVersion & systemCurrent BenchLM aggregated public evaluation | BenchLM public benchmark leaderboardsObserved Checked |
From result to context
Harvard-MIT Mathematics Tournament February 2026 problems used as a contest math evaluation.
The summary shows one best compatible source result per canonical product. Versions, effort settings and execution systems remain attached to the underlying records.
Scores from different versions or harnesses may not be interchangeable. The published source rows preserve those distinctions and their original units. A source result is not automatically an input to a current index.
This catalogue entry preserves available source evidence. Current indices admit only their specifically reviewed tracks and configurations.
Contamination risk: Unknown. Lifecycle: Active.