Reasoning
Mathematical and scientific problem solving.
Model profile / Mistral AI
Mistral AI's dense 128B Mistral Medium 3.5 model with 256K context, text and image input, configurable reasoning, function calling, structured outputs and Modified MIT licensing.
Four areas of evidence
Mathematical and scientific problem solving.
Software engineering and programming.
Factual reliability and instruction following.
Tool use, workflows and professional tasks.
Current index snapshot · 2026-09-16. Each category uses its own scale; scores across categories are not directly comparable. Counts describe published catalogue evidence. A point lead does not establish superiority.
Performance in context
Mistral Medium 3.5 128B has no current Capabilities rating. Explore the rated models below.
Lumina Capabilities Index
The current Capabilities cohort, with available operational measurements.
Chart loads as you explore
Lumina Capabilities Index
The current Capabilities cohort, with available operational measurements.
Chart loads as you explore
Explore the detail
No current index estimate. Available benchmark measurements and source records are retained below.
| Benchmark | Result | Published configuration | Source |
|---|---|---|---|
| AA Long Context Reasoningstandard | 65.333% | Mistral Medium 3.5 (Artificial Analysis completed independent run)Artificial Analysis AA-LCR evaluation | Mistral Medium 3.5 current model and individual-evaluation page ↗Observed 2026-08-29 · checked 2026-08-29independently-verified · ranking-eligible |
| CritPtstandard | 0% | Mistral Medium 3.5 (Artificial Analysis completed independent run)Artificial Analysis CritPt evaluation | Mistral Medium 3.5 current model and individual-evaluation page ↗Observed 2026-08-29 · checked 2026-08-29independently-verified · ranking-eligible |
| GDPval-AA v2v2 | 21.783% | Mistral Medium 3.5 (Artificial Analysis completed independent run)Artificial Analysis GDPval-AA v2 | Mistral Medium 3.5 current model and individual-evaluation page ↗Observed 2026-08-29 · checked 2026-08-29independently-verified · ranking-eligible |
| GPQA Diamonddiamond | 74.848% | Mistral Medium 3.5 (Artificial Analysis completed independent run)Artificial Analysis GPQA Diamond evaluation | Mistral Medium 3.5 current model and individual-evaluation page ↗Observed 2026-08-29 · checked 2026-08-29independently-verified · ranking-eligible |
| Humanity's Last Examtext-only current | 13.763% | Mistral Medium 3.5 (Artificial Analysis completed independent run)Artificial Analysis text-only HLE evaluation | Mistral Medium 3.5 current model and individual-evaluation page ↗Observed 2026-08-29 · checked 2026-08-29independently-verified · ranking-eligible |
| SciCode2024 | 39.583% | Mistral Medium 3.5 (Artificial Analysis completed independent run)Artificial Analysis SciCode evaluation | Mistral Medium 3.5 current model and individual-evaluation page ↗Observed 2026-08-29 · checked 2026-08-29independently-verified · reference-only |
| τ³-Banking3 | 15.052% | Mistral Medium 3.5 (Artificial Analysis completed independent run)Artificial Analysis tau3-Banking evaluation | Mistral Medium 3.5 current model and individual-evaluation page ↗Observed 2026-08-29 · checked 2026-08-29independently-verified · ranking-eligible |
| Terminal-Bench2.1 | 50.562% | Mistral Medium 3.5 (Artificial Analysis completed independent run)Artificial Analysis Terminal-Bench v2.1 in e2b | Mistral Medium 3.5 current model and individual-evaluation page ↗Observed 2026-08-29 · checked 2026-08-29independently-verified · ranking-eligible |
| Artificial Analysis Agentic Index2026 | 19 index | Exact BenchLM registry variant Mistral Medium 3.5 128B; bulk export does not retain a complete upstream harness configuration.Source-native system | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21source-checked · reference-only |
| Artificial Analysis Agentic Index2026 | 19 index | Exact BenchLM registry variant Mistral Medium 3.5 128B; bulk export does not retain a complete upstream harness configuration.Source-native system | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27source-checked · reference-only |