Reasoning
Mathematical and scientific problem solving.
Score and research range
- Full-precision point
- 102.42445862815174
- Range
- 99.81–105.04 · 90% Research Range
- Scope
- Conditional fuller R0 A
Score includes qualified prediction of missing evidence.
Model profile / Mistral AI
Mistral AI's open-weight multimodal model for agentic and coding use cases.
Four areas of evidence
Mathematical and scientific problem solving.
Score includes qualified prediction of missing evidence.
Software engineering and programming.
Score uses the admitted direct native evidence under the selected method.
Factual reliability and instruction following.
Tool use, workflows and professional tasks.
Score uses the admitted direct native evidence under the selected method.
Current index snapshot · 2026-09-16. Each category uses its own scale; scores across categories are not directly comparable. Evidence counts describe the retained Capabilities inputs; a family can contain several tasks or configurations. A point lead does not establish superiority.
Performance in context
Mistral Medium 3.5 is highlighted wherever a compatible measurement is available.
Lumina Capabilities Index
The current Capabilities cohort, with this model highlighted.
Chart loads as you explore
Lumina Capabilities Index
The current Capabilities cohort, with this model highlighted.
Chart loads as you explore
Explore the detail
Score uses the admitted direct native evidence under the selected method.
Score uses the admitted direct native evidence under the selected method.
Score includes qualified prediction of missing evidence.
Score uses the admitted direct native evidence under the selected method.
| Benchmark | Result | Published configuration | Source |
|---|---|---|---|
| Humanity's Last ExamSource-native version | 12.79% | Artificial Analysis public model evaluation page; exact provider variant as listed on the page.Artificial Analysis evaluation harness | Artificial Analysis evaluations for mistral-medium-3-5 ↗Observed 2026-07-15 · checked 2026-07-15source-checked · reference-only |
| GPQA DiamondSource-native version | 74.85% | Artificial Analysis public model evaluation page; exact provider variant as listed on the page.Artificial Analysis evaluation harness | Artificial Analysis evaluations for mistral-medium-3-5 ↗Observed 2026-07-15 · checked 2026-07-15source-checked · reference-only |
| MMMU-ProSource-native version | 64.86% | Artificial Analysis public model evaluation page; exact provider variant as listed on the page.Artificial Analysis evaluation harness | Artificial Analysis evaluations for mistral-medium-3-5 ↗Observed 2026-07-15 · checked 2026-07-15source-checked · ranking-eligible |
| Terminal-Bench2.1 | 50.56% | Artificial Analysis public model evaluation page; exact provider variant as listed on the page.Artificial Analysis evaluation harness | Artificial Analysis evaluations for mistral-medium-3-5 ↗Observed 2026-07-15 · checked 2026-07-15source-checked · reference-only |
| GDPval-AA v2Source-native version | 21.448% | Mistral Medium 3.5Artificial Analysis independent evaluation | GDPval-AA v2 Leaderboard ↗Observed 2026-07-15 · checked 2026-07-15source-checked · ranking-eligible |
| AA Long Context ReasoningSource-native version | 61% | Mistral Medium 3.5Artificial Analysis independent evaluation | AA Long Context Reasoning Leaderboard ↗Observed 2026-07-15 · checked 2026-07-15source-checked · reference-only |
| AA-OmniscienceSource-native version | -36.317 index | Mistral Medium 3.5Artificial Analysis independent evaluation | AA-Omniscience Leaderboard ↗Observed 2026-07-15 · checked 2026-07-15source-checked · reference-only |
| CritPtSource-native version | 0% | Mistral Medium 3.5Artificial Analysis independent evaluation | CritPt Leaderboard ↗Observed 2026-07-15 · checked 2026-07-15source-checked · ranking-eligible |
| τ²-Bench TelecomSource-native version | 94.152% | Mistral Medium 3.5Artificial Analysis independent evaluation | τ²-Bench Telecom Leaderboard ↗Observed 2026-07-15 · checked 2026-07-15source-checked · reference-only |
| τ³-BankingSource-native version | 14.433% | Mistral Medium 3.5Artificial Analysis independent evaluation | τ³-Banking Leaderboard ↗Observed 2026-07-15 · checked 2026-07-15source-checked · reference-only |