Reasoning
Mathematical and scientific problem solving.
Score and research range
- Full-precision point
- 121.99231190933563
- Range
- 118.61–124.98 · 90% Research Range
- Scope
- Conditional fuller R0 X
Score includes qualified prediction of missing evidence.
Model profile / Anthropic
Anthropic's active Fable-class model for demanding reasoning and long-horizon agentic work, sharing one underlying model with the restricted-access Claude Mythos 5.1 safeguard variant.
Four areas of evidence
Mathematical and scientific problem solving.
Score includes qualified prediction of missing evidence.
Software engineering and programming.
Score uses the admitted direct native evidence under the selected method.
Factual reliability and instruction following.
Tool use, workflows and professional tasks.
Current index snapshot · 2026-09-16. Each category uses its own scale; scores across categories are not directly comparable. Evidence counts describe the retained Capabilities inputs; a family can contain several tasks or configurations. A point lead does not establish superiority.
Performance in context
Claude Fable 5.1 is highlighted wherever a compatible measurement is available.
Lumina Capabilities Index
The current Capabilities cohort, with this model highlighted.
Chart loads as you explore
Lumina Capabilities Index
The current Capabilities cohort, with this model highlighted.
Chart loads as you explore
Explore the detail
Score uses the admitted direct native evidence under the selected method.
Score uses the admitted direct native evidence under the selected method.
Score includes qualified prediction of missing evidence.
| Benchmark | Result | Published configuration | Source |
|---|---|---|---|
| CursorBench3.2.0 | 73.4% | Claude Fable 5.1 Max on CursorBench 3.2.0system:cursorbench-3-2:claude-fable-5-1-max | CursorBench 3.2.0 cost savings benchmark table ↗Observed 2026-09-01 · checked 2026-09-01source-checked · reference-only |
| Benchmark | Result | Execution & attribution | Original source |
|---|---|---|---|
| FrontierCode 1.1 MainFrontierCode 1.1 Main · Cognition release table | 50.9percent (published) | Release-table configuration; per-model effort and harness vary. See original footnotes.provider measured competitorSource-native comparison only. Different settings are not silently substituted into retained index protocols. | Cognition ↗Published 2026-09-10 · checked 2026-09-16 |
| Terminal-Bench 4Terminal-Bench 4 · Cognition release table | 55.8percent (published) | Release-table configuration; per-model effort and harness vary. See original footnotes.provider measured competitorSource-native comparison only. Different settings are not silently substituted into retained index protocols. | Cognition ↗Published 2026-09-10 · checked 2026-09-16 |
| Terminal-Bench 2.1Terminal-Bench 2.1 · Cognition release table | 91.4percent (published) | Release-table configuration; per-model effort and harness vary. See original footnotes.provider measured competitorSource-native comparison only. Different settings are not silently substituted into retained index protocols. | Cognition ↗Published 2026-09-10 · checked 2026-09-16 |
| DeepSWE 1.1DeepSWE 1.1 · Cognition release table | 67.4percent (published) | Release-table configuration; per-model effort and harness vary. See original footnotes.provider measured competitorSource-native comparison only. Different settings are not silently substituted into retained index protocols. | Cognition ↗Published 2026-09-10 · checked 2026-09-16 |
No specialist task observations retained for this model.