A task with its own rules.
American Invitational Mathematics Examination 2025 contest problems used as a frontier math evaluation.
- Organisation
- MAA / public contest evals
- Version
- Current / rolling
Mathematics / Benchmark profile
American Invitational Mathematics Examination 2025 contest problems used as a frontier math evaluation.
Observed results
20 ranked models · higher is better · labels show rank and score
| Model | Rank | Model ID | Provider | Score | Unit |
|---|---|---|---|---|---|
| GPT-5.2 | #1 | gpt-5-2 | OpenAI | 99 | percent |
| GPT-5 Codex | #2 | gpt-5-codex | OpenAI | 98.6667 | percent |
| MAI-Thinking-1 | #3 | mai-thinking-1 | Microsoft | 97 | percent |
| Kimi K2.5 | #4 | kimi-k2-5 | Moonshot AI | 96.1 | percent |
| GLM-4.7 | #5 | glm-4-7 | Z.AI | 95.7 | percent |
| Kimi K2 Thinking | #6 | kimi-k2-thinking | Moonshot AI | 94.6667 | percent |
| GPT-5 | #7 | gpt-5 | OpenAI | 94.3333 | percent |
| Grok 4 | #8 | grok-4 | xAI | 92.6667 | percent |
| DeepSeek V3.2 | #9 | deepseek-v3-2 | DeepSeek | 92 | percent |
| o4-mini | #10 | o4-mini | OpenAI | 90.6667 | percent |
| DeepSeek V3.1 Terminus | #11 | deepseek-v3-1-terminus | DeepSeek | 89.6667 | percent |
| Grok 4 Fast | #12 | grok-4-fast | xAI | 89.6667 | percent |
| Nova 2 Pro | #13 | nova-2-pro | Amazon | 89 | percent |
| o3 | #14 | o3 | OpenAI | 88.3333 | percent |
| Claude Sonnet 4.5 | #15 | claude-sonnet-4-5 | Anthropic | 88 | percent |
| Gemini 2.5 Pro | #16 | gemini-2-5-pro | 87.6667 | percent | |
| GLM-4.6 | #17 | glm-4-6 | Z.AI | 86 | percent |
| Exaone 4.0 32B | #18 | exaone-4-0-32b | LG AI Research | 85.3 | percent |
| GPT-5 mini | #19 | gpt-5-mini | OpenAI | 85 | percent |
| Grok 3 mini | #20 | grok-3-mini | xAI | 84.6667 | percent |
| MiniMax M2.1 | #21 | minimax-m2-1 | MiniMax | 82.6667 | percent |
| Nemotron 3 Nano Omni 30B A3B | #22 | nemotron-3-nano-omni-30b-a3b | NVIDIA | 82.1 | percent |
| Claude Opus 4.1 | #23 | claude-opus-4-1 | Anthropic | 80.3333 | percent |
| Gemini 2.5 Flash | #24 | gemini-2-5-flash | 78.3333 | percent | |
| MiniMax M2 | #25 | minimax-m2 | MiniMax | 78.3333 | percent |
| Claude Sonnet 4 | #26 | claude-sonnet-4 | Anthropic | 74.3333 | percent |
| Claude Opus 4 | #27 | claude-opus-4 | Anthropic | 73.3333 | percent |
| Kimi K2 0905 | #28 | kimi-k2-0905 | Moonshot AI | 57.3333 | percent |
| Claude Sonnet 3.7 | #29 | claude-sonnet-3-7 | Anthropic | 56.3333 | percent |
| Grok Code Fast 1 | #30 | grok-code-fast-1 | xAI | 43.3333 | percent |
| LFM2.5-8B-A1B | #31 | lfm2-5-8b-a1b | LiquidAI | 42.53 | percent |
| MiniCPM5-1B | #32 | minicpm5-1b | OpenBMB | 40.42 | percent |
One best compatible score per canonical product · higher is better
| Rank | Model | Weights | Evidence | Score |
|---|---|---|---|---|
| 01 | GPT-5.2OpenAI | Closed weights | Public reference | 99% |
| 02 | GPT-5 CodexOpenAI | Closed weights | Public reference | 98.7% |
| 03 | MAI-Thinking-1Microsoft | Not documented | Source carrier | 97% |
| 04 | Kimi K2.5Moonshot AI | Open weights | Source carrier | 96.1% |
| 05 | GLM-4.7Z.AI | Open weights | Source carrier | 95.7% |
| 06 | Kimi K2 ThinkingMoonshot AI | Open weights | Public reference | 94.7% |
| 07 | GPT-5OpenAI | Closed weights | Public reference | 94.3% |
| 08 | Grok 4xAI | Closed weights | Public reference | 92.7% |
| 09 | DeepSeek V3.2DeepSeek | Open weights | Public reference | 92% |
| 10 | o4-miniOpenAI | Closed weights | Public reference | 90.7% |
Each distinct published measurement is retained. Same-snapshot canonical and configuration projections appear once.
| Published model / configuration | Score | Evidence & protocol | Source & dates |
|---|---|---|---|
Exact BenchLM registry variant MAI-Thinking-1; bulk export does not retain a complete upstream harness configuration.MAI-Thinking-1Exact identitySource label without a registered configuration ID Canonical product: mai-thinking-1 | 97% | Source carriersource-checkedVersion & system2025 Source-native system | BenchLM public datasets — 2026-08-01Observed Checked |
Exact BenchLM registry variant Kimi K2.5; bulk export does not retain a complete upstream harness configuration.Kimi K2.5Exact identitySource label without a registered configuration ID Canonical product: kimi-k2-5 | 96.1% | Source carriersource-checkedVersion & system2025 Source-native system | BenchLM public datasets — 2026-08-01Observed Checked |
Exact BenchLM registry variant Kimi K2.5 (Reasoning); bulk export does not retain a complete upstream harness configuration.Kimi K2.5Exact identitykimi-k2-5-thinking Canonical product: kimi-k2-5 | 96.1% | Source carriersource-checkedVersion & system2025 Source-native system | BenchLM public datasets — 2026-08-01Observed Checked |
Exact BenchLM registry variant GLM-4.7; bulk export does not retain a complete upstream harness configuration.GLM-4.7Exact identitySource label without a registered configuration ID Canonical product: glm-4-7 | 95.7% | Source carriersource-checkedVersion & system2025 Source-native system | BenchLM public datasets — 2026-08-01Observed Checked |
Exact BenchLM registry variant Claude Sonnet 4.5; bulk export does not retain a complete upstream harness configuration.Claude Sonnet 4.5Exact identitySource label without a registered configuration ID Canonical product: claude-sonnet-4-5 | 87% | Source carriersource-checkedVersion & system2025 Source-native system | BenchLM public datasets — 2026-08-01Observed Checked |
Exact BenchLM registry variant Exaone 4.0 32B; bulk export does not retain a complete upstream harness configuration.Exaone 4.0 32BExact identitySource label without a registered configuration ID Canonical product: exaone-4-0-32b | 85.3% | Source carriersource-checkedVersion & system2025 Source-native system | BenchLM public datasets — 2026-08-01Observed Checked |
Exact BenchLM registry variant Nemotron 3 Nano Omni 30B A3B; bulk export does not retain a complete upstream harness configuration.Nemotron 3 Nano Omni 30B A3BExact identitySource label without a registered configuration ID Canonical product: nemotron-3-nano-omni-30b-a3b | 82.1% | Source carriersource-checkedVersion & system2025 Source-native system | BenchLM public datasets — 2026-08-01Observed Checked |
Exact BenchLM registry variant LFM2.5-8B-A1B; bulk export does not retain a complete upstream harness configuration.LFM2.5-8B-A1BExact identitySource label without a registered configuration ID Canonical product: lfm2-5-8b-a1b | 42.5% | Source carriersource-checkedVersion & system2025 Source-native system | BenchLM public datasets — 2026-08-01Observed Checked |
Exact BenchLM registry variant MiniCPM5-1B; bulk export does not retain a complete upstream harness configuration.MiniCPM5-1BExact identitySource label without a registered configuration ID Canonical product: minicpm5-1b | 40.4% | Source carriersource-checkedVersion & system2025 Source-native system | BenchLM public datasets — 2026-08-01Observed Checked |
Exact BenchLM registry variant MAI-Thinking-1; bulk export does not retain a complete upstream harness configuration.MAI-Thinking-1Exact identitySource label without a registered configuration ID Canonical product: mai-thinking-1 | 97% | Source carriersource-checkedVersion & system2025 Source-native system | BenchLM public datasets — 2026-07-27Observed Checked |
From result to context
American Invitational Mathematics Examination 2025 contest problems used as a frontier math evaluation.
The summary shows one best compatible source result per canonical product. Versions, effort settings and execution systems remain attached to the underlying records.
Scores from different versions or harnesses may not be interchangeable. The published source rows preserve those distinctions and their original units. A source result is not automatically an input to a current index.
This catalogue entry preserves available source evidence. Current indices admit only their specifically reviewed tracks and configurations.
Contamination risk: Unknown. Lifecycle: Active.