A task with its own rules.
Cursor's first-party benchmark for ambiguous multi-file coding-agent tasks from real Cursor sessions (v3.2 public snapshot).
- Organisation
- Cursor
- Version
- 3.2
Coding / Benchmark profile
Cursor's first-party benchmark for ambiguous multi-file coding-agent tasks from real Cursor sessions (v3.2 public snapshot).
Observed results
12 ranked models · higher is better · labels show rank and score
| Model | Rank | Model ID | Provider | Score | Unit |
|---|---|---|---|---|---|
| Claude Fable 5.1 | #1 | claude-fable-5-1 | Anthropic | 73.4 | percent |
| Claude Fable 5 | #2 | claude-fable-5 | Anthropic | 70.5 | percent |
| Grok 4.6 | #3 | grok-4-6 | xAI | 69.9 | percent |
| GPT-5.6 Sol | #4 | gpt-5-6-sol | OpenAI | 67.2 | percent |
| Grok 4.5 | #5 | grok-4-5 | xAI | 66.7 | percent |
| GPT-5.6 Terra | #6 | gpt-5-6-terra | OpenAI | 64.9 | percent |
| Claude Opus 4.8 | #7 | claude-opus-4-8 | Anthropic | 62.3 | percent |
| Claude Sonnet 5 | #8 | claude-sonnet-5 | Anthropic | 61.5 | percent |
| GPT-5.6 Luna | #9 | gpt-5-6-luna | OpenAI | 61.1 | percent |
| GPT-5.5 | #10 | gpt-5-5 | OpenAI | 58.4 | percent |
| GLM-5.2 | #11 | glm-5-2 | Z.AI | 55 | percent |
| Gemini 3.5 Flash | #12 | gemini-3-5-flash | 48.8 | percent |
One best compatible score per canonical product · higher is better
| Rank | Model | Weights | Evidence | Score |
|---|---|---|---|---|
| 01 | Claude Fable 5.1Anthropic | Closed weights | Public reference | 73.4% |
| 02 | Claude Fable 5Anthropic | Closed weights | Source carrier | 70.5% |
| 03 | Grok 4.6xAI | Closed weights | Public reference | 69.9% |
| 04 | GPT-5.6 SolOpenAI | Closed weights | Source carrier | 67.2% |
| 05 | Grok 4.5xAI | Closed weights | Source carrier | 66.7% |
| 06 | GPT-5.6 TerraOpenAI | Closed weights | Source carrier | 64.9% |
| 07 | Claude Opus 4.8Anthropic | Closed weights | Source carrier | 62.3% |
| 08 | Claude Sonnet 5Anthropic | Closed weights | Source carrier | 61.5% |
| 09 | GPT-5.6 LunaOpenAI | Closed weights | Source carrier | 61.1% |
| 10 | GPT-5.5OpenAI | Closed weights | Source carrier | 58.4% |
Each distinct published measurement is retained. Same-snapshot canonical and configuration projections appear once.
| Published model / configuration | Score | Evidence & protocol | Source & dates |
|---|---|---|---|
Claude Fable 5.1 Max on CursorBench 3.2.0Claude Fable 5.1Exact identityclaude-fable-5-1-max Canonical product: claude-fable-5-1 | 73.4% | Public referencesource-checkedVersion & system3.2.0 system:cursorbench-3-2:claude-fable-5-1-max | CursorBench 3.2.0 cost savings benchmark tableObserved Checked |
Fable 5 MaxClaude Fable 5Exact identityclaude-fable-5-max Canonical product: claude-fable-5 | 70.5% | Public referenceprovider-reportedVersion & system3.2 xAI Grok 4.6 official release comparison table | Grok 4.6Observed Checked |
Grok 4.6 HighGrok 4.6Exact identitygrok-4-6-high Canonical product: grok-4-6 | 69.9% | Public referenceprovider-reportedVersion & system3.2 xAI Grok 4.6 official release comparison table | Grok 4.6Observed Checked |
GPT-5.6 Sol MaxGPT-5.6 SolExact identitygpt-5-6-sol-max Canonical product: gpt-5-6-sol | 67.2% | Public referenceprovider-reportedVersion & system3.2 xAI Grok 4.6 official release comparison table | Grok 4.6Observed Checked |
Grok 4.5 HighGrok 4.5Exact identitygrok-4-5-aa-2-high Canonical product: grok-4-5 | 66.7% | Public referenceprovider-reportedVersion & system3.2 xAI Grok 4.6 official release comparison table | Grok 4.6Observed Checked |
BenchLM public leaderboard / CursorBench snapshot for the exact model variant.Claude Fable 5Exact identitySource label without a registered configuration ID Canonical product: claude-fable-5 | 70.5% | Source carriersource-checkedVersion & system3.2 BenchLM / Cursor public evaluation | CursorBench public evaluation via BenchLMObserved Checked |
Exact model variant as listed on the BenchLM public benchmark leaderboard page.Claude Fable 5Exact identitySource label without a registered configuration ID Canonical product: claude-fable-5 | 70.5% | Source carriersource-checkedVersion & system3.2 BenchLM aggregated public evaluation | BenchLM public benchmark leaderboardsObserved Checked |
BenchLM public leaderboard / CursorBench snapshot for the exact model variant.GPT-5.6 SolExact identitySource label without a registered configuration ID Canonical product: gpt-5-6-sol | 67.2% | Source carriersource-checkedVersion & system3.2 BenchLM / Cursor public evaluation | CursorBench public evaluation via BenchLMObserved Checked |
Exact model variant as listed on the BenchLM public benchmark leaderboard page.GPT-5.6 SolExact identitySource label without a registered configuration ID Canonical product: gpt-5-6-sol | 67.2% | Source carriersource-checkedVersion & system3.2 BenchLM aggregated public evaluation | BenchLM public benchmark leaderboardsObserved Checked |
BenchLM public leaderboard / CursorBench snapshot for the exact model variant.Grok 4.5Exact identitySource label without a registered configuration ID Canonical product: grok-4-5 | 66.7% | Source carriersource-checkedVersion & system3.2 BenchLM / Cursor public evaluation | CursorBench public evaluation via BenchLMObserved Checked |
From result to context
Cursor's first-party benchmark for ambiguous multi-file coding-agent tasks from real Cursor sessions (v3.2 public snapshot).
The summary shows one best compatible source result per canonical product. Versions, effort settings and execution systems remain attached to the underlying records.
Scores from different versions or harnesses may not be interchangeable. The published source rows preserve those distinctions and their original units. A source result is not automatically an input to a current index.
A reviewed track from this benchmark is represented in the following current indices. Each index applies its own protocol and eligibility rules.
Contamination risk: Medium. Lifecycle: Active.