A task with its own rules.
American Invitational Mathematics Examination 2026 contest problems used as a frontier math evaluation.
- Organisation
- MAA / public contest evals
- Version
- Current / rolling
Mathematics / Benchmark profile
American Invitational Mathematics Examination 2026 contest problems used as a frontier math evaluation.
Observed results
20 ranked models · higher is better · labels show rank and score
| Model | Rank | Model ID | Provider | Score | Unit |
|---|---|---|---|---|---|
| GLM-5.2 | #1 | glm-5-2 | Z.AI | 99.2 | percent |
| Inkling | #2 | inkling | Thinking Machines Lab | 97.1 | percent |
| Kimi K2.6 | #3 | kimi-k2-6 | Moonshot AI | 96.4 | percent |
| GLM-5 | #4 | glm-5 | Z.AI | 95.8 | percent |
| Kimi K2.5 | #5 | kimi-k2-5 | Moonshot AI | 95.8 | percent |
| Solar Open 2 250B | #6 | solar-open2-250b | Upstage | 95.7 | percent |
| Inkling-Small | #7 | inkling-small | Thinking Machines Lab | 95.5 | percent |
| GLM-5.1 | #8 | glm-5-1 | Z.AI | 95.3 | percent |
| Qwen3.6 Plus | #9 | qwen3-6-plus | Alibaba Cloud | 95.3 | percent |
| Claude Opus 4.5 | #10 | claude-opus-4-5 | Anthropic | 95.1 | percent |
| Muse Glimmer 30B | #11 | muse-glimmer-30b | Meta | 94.7 | percent |
| MAI-Thinking-1 | #12 | mai-thinking-1 | Microsoft | 94.5 | percent |
| Qwen3.6 27B | #13 | qwen3-6-27b | Alibaba Cloud | 94.1 | percent |
| Qwen3.5 397B A17B | #14 | qwen3-5-397b | Alibaba Cloud | 93.3 | percent |
| Qwen3.6-35B-A3B | #15 | qwen3-6-35b-a3b | Alibaba Cloud | 92.7 | percent |
| ZAYA1-8B | #16 | zaya1-8b | Zyphra | 89.1 | percent |
| Gemma 4 12B Unified | #17 | gemma-4-12b | 77.5 | percent | |
| ZAYA1-74B-Preview | #18 | zaya1-74b-preview | Zyphra | 76.4 | percent |
| LFM2.5-8B-A1B | #19 | lfm2-5-8b-a1b | LiquidAI | 50 | percent |
| MiniCPM5-1B | #20 | minicpm5-1b | OpenBMB | 40.42 | percent |
One best compatible score per canonical product · higher is better
| Rank | Model | Weights | Evidence | Score |
|---|---|---|---|---|
| 01 | GLM-5.2Z.AI | Open weights | Source carrier | 99.2% |
| 02 | InklingThinking Machines Lab | Open weights | Source carrier | 97.1% |
| 03 | Kimi K2.6Moonshot AI | Open weights | Source carrier | 96.4% |
| 04 | GLM-5Z.AI | Open weights | Source carrier | 95.8% |
| 05 | Kimi K2.5Moonshot AI | Open weights | Source carrier | 95.8% |
| 06 | Solar Open 2 250BUpstage | Open weights | Direct source result | 95.7% |
| 07 | Inkling-SmallThinking Machines Lab | Open weights | Source carrier | 95.5% |
| 08 | GLM-5.1Z.AI | Open weights | Source carrier | 95.3% |
| 09 | Qwen3.6 PlusAlibaba Cloud | Closed weights | Source carrier | 95.3% |
| 10 | Claude Opus 4.5Anthropic | Closed weights | Source carrier | 95.1% |
Each distinct published measurement is retained. Same-snapshot canonical and configuration projections appear once.
| Published model / configuration | Score | Evidence & protocol | Source & dates |
|---|---|---|---|
DeepSeek V4 Flash max as published by UpstageDeepSeek V4 FlashExact identitydeepseek-v4-flash-max Canonical product: deepseek-v4-flash | 97% | Relative comparisonprovider-reportedVersion & system2026 Solar Open 2 card English table | Solar Open 2 250B model cardObserved Checked |
Command A+ as published by UpstageCommand A+Exact identitycommand-a-plus-solar-open2-unspecified Canonical product: command-a-plus | 96% | Relative comparisonprovider-reportedVersion & system2026 Solar Open 2 card English table | Solar Open 2 250B model cardObserved Checked |
Solar Open 2 250B (high)Solar Open 2 250BExact identitysolar-open2-250b-high Canonical product: solar-open2-250b | 95.7% | Direct source resultprovider-reportedVersion & system2026 Solar Open 2 card English table | Solar Open 2 250B model cardObserved Checked |
MiMo-V2.5 as published by UpstageMiMo-V2.5Exact identitymimo-v2-5-solar-open2-unspecified Canonical product: mimo-v2-5 | 92.3% | Relative comparisonprovider-reportedVersion & system2026 Solar Open 2 card English table | Solar Open 2 250B model cardObserved Checked |
Mistral Medium 3.5 as published by UpstageMistral Medium 3.5Exact identitymistral-medium-3-5-high Canonical product: mistral-medium-3-5 | 89% | Relative comparisonprovider-reportedVersion & system2026 Solar Open 2 card English table | Solar Open 2 250B model cardObserved Checked |
Solar Open 100B as published by UpstageSolar Open 100B (Reasoning)Exact identitysolar-open-100b-reasoning-high Canonical product: solar-open-100b-reasoning | 87.7% | Relative comparisonprovider-reportedVersion & system2026 Solar Open 2 card English table | Solar Open 2 250B model cardObserved Checked |
Muse Glimmer-30B; high reasoning; temperature=1.0; top_p=0.95; top_k=64Muse Glimmer 30BExact identitymuse-glimmer-30b-high Canonical product: muse-glimmer-30b | 94.7% | Public referenceprovider-reportedVersion & system2026 / 30 questions Meta AIME 2026 evaluation averaged across ten runs | Muse Glimmer Evaluation MethodologyObserved Checked |
Exact BenchLM registry variant GLM-5.2; bulk export does not retain a complete upstream harness configuration.GLM-5.2Exact identitySource label without a registered configuration ID Canonical product: glm-5-2 | 99.2% | Source carriersource-checkedVersion & system2026 Source-native system | BenchLM public datasets — 2026-08-01Observed Checked |
Exact BenchLM registry variant Inkling; bulk export does not retain a complete upstream harness configuration.InklingExact identitySource label without a registered configuration ID Canonical product: inkling | 97.1% | Source carriersource-checkedVersion & system2026 Source-native system | BenchLM public datasets — 2026-08-01Observed Checked |
Exact BenchLM registry variant Kimi K2.6; bulk export does not retain a complete upstream harness configuration.Kimi K2.6Exact identitySource label without a registered configuration ID Canonical product: kimi-k2-6 | 96.4% | Source carriersource-checkedVersion & system2026 Source-native system | BenchLM public datasets — 2026-08-01Observed Checked |
From result to context
American Invitational Mathematics Examination 2026 contest problems used as a frontier math evaluation.
The summary shows one best compatible source result per canonical product. Versions, effort settings and execution systems remain attached to the underlying records.
Scores from different versions or harnesses may not be interchangeable. The published source rows preserve those distinctions and their original units. A source result is not automatically an input to a current index.
This catalogue entry preserves available source evidence. Current indices admit only their specifically reviewed tracks and configurations.
Contamination risk: Unknown. Lifecycle: Active.