Reasoning
Mathematical and scientific problem solving.
Score and research range
- Full-precision point
- 108.75700750527959
- Range
- 106.14–111.38 · 90% Research Range
- Scope
- Conditional fuller R0 A
Score includes qualified prediction of missing evidence.
Model profile / Xiaomi
MiMo-V2.5 is a open-weight model tracked via Artificial Analysis independent evaluations for capability comparison across coding, agents, and reasoning.
Four areas of evidence
Mathematical and scientific problem solving.
Score includes qualified prediction of missing evidence.
Software engineering and programming.
Factual reliability and instruction following.
Tool use, workflows and professional tasks.
Score uses the admitted direct native evidence under the selected method.
Current index snapshot · 2026-09-16. Each category uses its own scale; scores across categories are not directly comparable. Evidence counts describe the retained Capabilities inputs; a family can contain several tasks or configurations. A point lead does not establish superiority.
Performance in context
MiMo-V2.5 is highlighted wherever a compatible measurement is available.
Lumina Capabilities Index
The current Capabilities cohort, with this model highlighted.
Chart loads as you explore
MiMo-V2.5: No exact base-profile release and documented 10K output measurement; other effort/protocol/alias records not substituted
Lumina Capabilities Index
The current Capabilities cohort, with this model highlighted.
Chart loads as you explore
Explore the detail
Score uses the admitted direct native evidence under the selected method.
Score includes qualified prediction of missing evidence.
Score uses the admitted direct native evidence under the selected method.
| Benchmark | Result | Published configuration | Source |
|---|---|---|---|
| CharXiv2024 | 81% | Exact BenchLM registry variant MiMo-V2.5; bulk export does not retain a complete upstream harness configuration.Source-native system | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21source-checked · reference-only |
| CharXiv2024 | 81% | Exact BenchLM registry variant MiMo-V2.5; bulk export does not retain a complete upstream harness configuration.Source-native system | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27source-checked · reference-only |
| CharXiv2024 | 81% | Exact BenchLM registry variant MiMo-V2.5; bulk export does not retain a complete upstream harness configuration.Source-native system | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01source-checked · reference-only |
| Claw-Eval2026 | 62.3% | Exact BenchLM registry variant MiMo-V2.5; bulk export does not retain a complete upstream harness configuration.Source-native system | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21source-checked · reference-only |
| Claw-Eval2026 | 62.3% | Exact BenchLM registry variant MiMo-V2.5; bulk export does not retain a complete upstream harness configuration.Source-native system | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27source-checked · reference-only |
| Claw-Eval2026 | 62.3% | Exact BenchLM registry variant MiMo-V2.5; bulk export does not retain a complete upstream harness configuration.Source-native system | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01source-checked · reference-only |
| Design Arena Website Elo2026 | 1,291 elo | Exact BenchLM registry variant MiMo-V2.5; bulk export does not retain a complete upstream harness configuration.Source-native system | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21source-checked · reference-only |
| Design Arena Website Elo2026 | 1,291 elo | Exact BenchLM registry variant MiMo-V2.5; bulk export does not retain a complete upstream harness configuration.Source-native system | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27source-checked · reference-only |
| Design Arena Website Elo2026 | 1,295 elo | Exact BenchLM registry variant MiMo-V2.5; bulk export does not retain a complete upstream harness configuration.Source-native system | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01source-checked · reference-only |
| Gert Labs Composite Game Benchmark2026 | 46.89% | Exact BenchLM registry variant MiMo-V2.5; bulk export does not retain a complete upstream harness configuration.Source-native system | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21source-checked · reference-only |