Reasoning
Mathematical and scientific problem solving.
Score and research range
- Full-precision point
- 103.2398650352577
- Range
- 100.62–105.86 · 90% Research Range
- Scope
- Conditional fuller R0 A
Score includes qualified prediction of missing evidence.
Model profile / DeepSeek
DeepSeek V3.1 is a open weight model variant recorded in the BenchLM public dataset.
Four areas of evidence
Mathematical and scientific problem solving.
Score includes qualified prediction of missing evidence.
Software engineering and programming.
Factual reliability and instruction following.
Tool use, workflows and professional tasks.
Current index snapshot · 2026-09-16. Each category uses its own scale; scores across categories are not directly comparable. Counts describe published catalogue evidence. A point lead does not establish superiority.
Performance in context
DeepSeek V3.1 has no current Capabilities rating. Explore the rated models below.
Lumina Capabilities Index
The current Capabilities cohort, with available operational measurements.
Chart loads as you explore
Lumina Capabilities Index
The current Capabilities cohort, with available operational measurements.
Chart loads as you explore
Explore the detail
Score includes qualified prediction of missing evidence.
| Benchmark | Result | Published configuration | Source |
|---|---|---|---|
| AA Long Context Reasoningstandard | 56.667% | DeepSeek V3.1 (Reasoning) (Artificial Analysis independent run)Artificial Analysis AA-LCR evaluation | DeepSeek V3.1 (Reasoning) individual evaluations ↗Observed 2026-08-29 · checked 2026-08-29independently-verified · reference-only |
| CritPtstandard | 2% | DeepSeek V3.1 (Reasoning) (Artificial Analysis independent run)Artificial Analysis CritPt evaluation | DeepSeek V3.1 (Reasoning) individual evaluations ↗Observed 2026-08-29 · checked 2026-08-29independently-verified · reference-only |
| GPQA Diamonddiamond | 77.879% | DeepSeek V3.1 (Reasoning) (Artificial Analysis independent run)Artificial Analysis GPQA Diamond evaluation | DeepSeek V3.1 (Reasoning) individual evaluations ↗Observed 2026-08-29 · checked 2026-08-29independently-verified · reference-only |
| Humanity's Last Examtext-only current | 14.272% | DeepSeek V3.1 (Reasoning) (Artificial Analysis independent run)Artificial Analysis text-only HLE evaluation | DeepSeek V3.1 (Reasoning) individual evaluations ↗Observed 2026-08-29 · checked 2026-08-29independently-verified · reference-only |
| SciCode2024 | 39.12% | DeepSeek V3.1 (Reasoning) (Artificial Analysis independent run)Artificial Analysis SciCode evaluation | DeepSeek V3.1 (Reasoning) individual evaluations ↗Observed 2026-08-29 · checked 2026-08-29independently-verified · reference-only |
| AA Long Context Reasoningstandard | 46.667% | DeepSeek V3.1 (Non-reasoning) (Artificial Analysis independent run)Artificial Analysis AA-LCR evaluation | DeepSeek V3.1 (Non-reasoning) individual evaluations ↗Observed 2026-08-29 · checked 2026-08-29independently-verified · ranking-eligible |
| CritPtstandard | 0% | DeepSeek V3.1 (Non-reasoning) (Artificial Analysis independent run)Artificial Analysis CritPt evaluation | DeepSeek V3.1 (Non-reasoning) individual evaluations ↗Observed 2026-08-29 · checked 2026-08-29independently-verified · ranking-eligible |
| GPQA Diamonddiamond | 73.535% | DeepSeek V3.1 (Non-reasoning) (Artificial Analysis independent run)Artificial Analysis GPQA Diamond evaluation | DeepSeek V3.1 (Non-reasoning) individual evaluations ↗Observed 2026-08-29 · checked 2026-08-29independently-verified · ranking-eligible |
| Humanity's Last Examtext-only current | 6.673% | DeepSeek V3.1 (Non-reasoning) (Artificial Analysis independent run)Artificial Analysis text-only HLE evaluation | DeepSeek V3.1 (Non-reasoning) individual evaluations ↗Observed 2026-08-29 · checked 2026-08-29independently-verified · ranking-eligible |
| SciCode2024 | 36.69% | DeepSeek V3.1 (Non-reasoning) (Artificial Analysis independent run)Artificial Analysis SciCode evaluation | DeepSeek V3.1 (Non-reasoning) individual evaluations ↗Observed 2026-08-29 · checked 2026-08-29independently-verified · reference-only |