A task with its own rules.
A provider-run Terminal-Bench 2.1 result stored separately from the repository's Terminal-Bench 2.0 lane.
- Organisation
- DeepSeek-AI
- Version
- 2026
Coding / Benchmark profile
A provider-run Terminal-Bench 2.1 result stored separately from the repository's Terminal-Bench 2.0 lane.
Observed results
1 ranked models · higher is better · labels show rank and score
| Model | Rank | Model ID | Provider | Score | Unit |
|---|---|---|---|---|---|
| DeepSeek V4 Flash | #1 | deepseek-v4-flash | DeepSeek | 82.7 | percent |
One best compatible score per canonical product · higher is better
| Rank | Model | Weights | Evidence | Score |
|---|---|---|---|---|
| 01 | DeepSeek V4 FlashDeepSeek | Open weights | Source carrier | 82.7% |
Each distinct published measurement is retained. Same-snapshot canonical and configuration projections appear once.
| Published model / configuration | Score | Evidence & protocol | Source & dates |
|---|---|---|---|
Exact BenchLM registry variant DeepSeek V4 Flash (Max); bulk export does not retain a complete upstream harness configuration.DeepSeek V4 FlashExact identitydeepseek-v4-flash-max Canonical product: deepseek-v4-flash | 82.7% | Source carriersource-checkedVersion & system2026 Source-native system | BenchLM public datasets — 2026-08-01Observed Checked |
| Model | Result | Execution & attribution | Original source |
|---|---|---|---|
| Grok 4.6Terminal-Bench 2.1 · Cognition release table | 88.4percent (published) | Release-table configuration; per-model effort and harness vary. See original footnotes.provider measured competitorSource-native comparison only. Different settings are not silently substituted into retained index protocols. | Cognition ↗Published 2026-09-10 · checked 2026-09-16 |
| Kimi K3Terminal-Bench 2.1 · Cognition release table | 88.3percent (published) | Release-table configuration; per-model effort and harness vary. See original footnotes.provider measured competitorSource-native comparison only. Different settings are not silently substituted into retained index protocols. | Cognition ↗Published 2026-09-10 · checked 2026-09-16 |
| GPT-6 AstraTerminal-Bench 2.1 · Cognition release table | 89.9percent (published) | Release-table configuration; per-model effort and harness vary. See original footnotes.provider measured competitorSource-native comparison only. Different settings are not silently substituted into retained index protocols. | Cognition ↗Published 2026-09-10 · checked 2026-09-16 |
| Fable 5.1Terminal-Bench 2.1 · Cognition release table | 91.4percent (published) | Release-table configuration; per-model effort and harness vary. See original footnotes.provider measured competitorSource-native comparison only. Different settings are not silently substituted into retained index protocols. | Cognition ↗Published 2026-09-10 · checked 2026-09-16 |
| SWE-1.7Terminal-Bench 2.1 · Cognition release table | 81.5percent (published) | Release-table configuration; per-model effort and harness vary. See original footnotes.provider subjectSource-native comparison only. Different settings are not silently substituted into retained index protocols. | Cognition ↗Published 2026-09-10 · checked 2026-09-16 |
| SWE-2Terminal-Bench 2.1 · Cognition release table | 92.8percent (published) | Release-table configuration; per-model effort and harness vary. See original footnotes.provider subjectSource-native comparison only. Different settings are not silently substituted into retained index protocols. | Cognition ↗Published 2026-09-10 · checked 2026-09-16 |
| GPT-5.6 SolTerminal-Bench 2.1 · Cognition release table | 88.8percent (published) | Release-table configuration; per-model effort and harness vary. See original footnotes.provider measured competitorSource-native comparison only. Different settings are not silently substituted into retained index protocols. | Cognition ↗Published 2026-09-10 · checked 2026-09-16 |
From result to context
A provider-run Terminal-Bench 2.1 result stored separately from the repository's Terminal-Bench 2.0 lane.
The summary shows one best compatible source result per canonical product. Versions, effort settings and execution systems remain attached to the underlying records.
Scores from different versions or harnesses may not be interchangeable. The published source rows preserve those distinctions and their original units. A source result is not automatically an input to a current index.
This catalogue entry preserves available source evidence. Current indices admit only their specifically reviewed tracks and configurations.
Contamination risk: Unknown. Lifecycle: Active.