Reasoning
Mathematical and scientific problem solving.
Score and research range
- Full-precision point
- 104.223754339087
- Range
- 101.60–106.84 · 90% Research Range
- Scope
- Conditional fuller R0 A
Score includes qualified prediction of missing evidence.
Model profile / DeepSeek
DeepSeek V3.1 Terminus is a open-weight model tracked via Artificial Analysis independent evaluations for capability comparison across coding, agents, and reasoning.
Four areas of evidence
Mathematical and scientific problem solving.
Score includes qualified prediction of missing evidence.
Software engineering and programming.
Factual reliability and instruction following.
Tool use, workflows and professional tasks.
Current index snapshot · 2026-09-16. Each category uses its own scale; scores across categories are not directly comparable. Evidence counts describe the retained Capabilities inputs; a family can contain several tasks or configurations. A point lead does not establish superiority.
Performance in context
DeepSeek V3.1 Terminus is highlighted wherever a compatible measurement is available.
DeepSeek V3.1 Terminus: Retired or changed served checkpoint; current DeepSeek endpoint prices cannot transfer to frozen predecessor weights
Lumina Capabilities Index
The current Capabilities cohort, with this model highlighted.
Chart loads as you explore
DeepSeek V3.1 Terminus: No exact base-profile release and documented 10K output measurement; other effort/protocol/alias records not substituted
Lumina Capabilities Index
The current Capabilities cohort, with this model highlighted.
Chart loads as you explore
Explore the detail
Score uses the admitted direct native evidence under the selected method.
Score includes qualified prediction of missing evidence.
| Benchmark | Result | Published configuration | Source |
|---|---|---|---|
| GDPval-AA v2Source-native version | 19.464% | DeepSeek V3.1 Terminus (Reasoning)Artificial Analysis independent evaluation | GDPval-AA v2 Leaderboard ↗Observed 2026-07-15 · checked 2026-07-15source-checked · ranking-eligible |
| AA Long Context ReasoningSource-native version | 65% | DeepSeek V3.1 Terminus (Reasoning)Artificial Analysis independent evaluation | AA Long Context Reasoning Leaderboard ↗Observed 2026-07-15 · checked 2026-07-15source-checked · reference-only |
| AA-OmniscienceSource-native version | -24.317 index | DeepSeek V3.1 Terminus (Reasoning)Artificial Analysis independent evaluation | AA-Omniscience Leaderboard ↗Observed 2026-07-15 · checked 2026-07-15source-checked · reference-only |
| CritPtSource-native version | 1.714% | DeepSeek V3.1 Terminus (Reasoning)Artificial Analysis independent evaluation | CritPt Leaderboard ↗Observed 2026-07-15 · checked 2026-07-15source-checked · ranking-eligible |
| τ²-Bench TelecomSource-native version | 37.135% | DeepSeek V3.1 Terminus (Reasoning)Artificial Analysis independent evaluation | τ²-Bench Telecom Leaderboard ↗Observed 2026-07-15 · checked 2026-07-15source-checked · reference-only |
| τ³-BankingSource-native version | 15.876% | DeepSeek V3.1 Terminus (Reasoning)Artificial Analysis independent evaluation | τ³-Banking Leaderboard ↗Observed 2026-07-15 · checked 2026-07-15source-checked · ranking-eligible |
| Terminal-Bench HardSource-native version | 30.303% | DeepSeek V3.1 Terminus (Reasoning)Artificial Analysis independent evaluation | Terminal-Bench Hard Leaderboard ↗Observed 2026-07-15 · checked 2026-07-15source-checked · reference-only |
| Terminal-BenchSource-native version | 44.944% | DeepSeek V3.1 Terminus (Reasoning)Artificial Analysis independent evaluation | Terminal-Bench Hard Leaderboard ↗Observed 2026-07-15 · checked 2026-07-15source-checked · ranking-eligible |
| SciCodeSource-native version | 40.625% | DeepSeek V3.1 Terminus (Reasoning)Artificial Analysis independent evaluation | Artificial Analysis Models Leaderboard ↗Observed 2026-07-15 · checked 2026-07-15source-checked · reference-only |
| Humanity's Last ExamSource-native version | 15.246% | DeepSeek V3.1 Terminus (Reasoning)Artificial Analysis independent evaluation | Artificial Analysis Models Leaderboard ↗Observed 2026-07-15 · checked 2026-07-15source-checked · reference-only |