A task with its own rules.
Contamination-resistant software engineering repair tasks across real repositories under a shared agent harness.
- Organisation
- DataCurve
- Version
- v1.1
Coding / Benchmark profile
Contamination-resistant software engineering repair tasks across real repositories under a shared agent harness.
Observed results
20 ranked models · higher is better · labels show rank and score
| Model | Rank | Model ID | Provider | Score | Unit |
|---|---|---|---|---|---|
| GPT-5.6 Sol | #1 | gpt-5-6-sol | OpenAI | 73 | percent |
| Claude Fable 5 | #2 | claude-fable-5 | Anthropic | 70 | percent |
| GPT-5.6 Terra | #3 | gpt-5-6-terra | OpenAI | 69.6 | percent |
| Claude Opus 5 | #4 | claude-opus-5 | Anthropic | 68.8 | percent |
| Kimi K3 | #5 | kimi-k3 | Moonshot AI | 67.5 | USD |
| GPT-5.6 Luna | #6 | gpt-5-6-luna | OpenAI | 67.2 | USD |
| GLM-5.3 | #7 | glm-5-3 | Z.AI | 66.9 | percent |
| Grok 4.6 | #8 | grok-4-6 | xAI | 65.9 | percent |
| Gemini 3.7 Flash | #9 | gemini-3-7-flash | 65.3 | percent | |
| DeepSeek V4 Pro 0813 | #10 | deepseek-v4-pro-0813 | DeepSeek | 62.7 | percent |
| DeepSeek V4 Flash Vision Exp | #11 | deepseek-v4-flash-vision-exp | DeepSeek | 59.3 | percent |
| Muse Spark 1.2 | #12 | muse-spark-1-2 | Meta | 59.3 | percent |
| Claude Opus 4.8 | #13 | claude-opus-4-8 | Anthropic | 58 | percent |
| Qwen3.8 Max | #14 | qwen-3-8-max | Alibaba Cloud | 56.6 | percent |
| DeepSeek V4 Flash | #15 | deepseek-v4-flash | DeepSeek | 54.4 | USD |
| DeepSeek V4 Flash 0731 | #16 | deepseek-v4-flash-0731 | DeepSeek | 54.4 | percent |
| Claude Sonnet 5 | #17 | claude-sonnet-5 | Anthropic | 54 | percent |
| Grok 4.5 | #18 | grok-4-5 | xAI | 54 | percent |
| Muse Spark 1.1 | #19 | muse-spark-1-1 | Meta | 53.3 | USD |
| Gemini 3.6 Flash | #20 | gemini-3-6-flash | 49 | percent | |
| GLM-5.2 | #21 | glm-5-2 | Z.AI | 46.2 | percent |
| Qwen3.8-27B | #22 | qwen-3-8-27b | Alibaba Cloud | 42.2 | percent |
| Laguna S 2.1 | #23 | laguna-s-2-1 | Poolside | 40.4 | USD |
| Gemini 3.5 Flash | #24 | gemini-3-5-flash | 37 | percent | |
| DeepSeek V4 Pro | #25 | deepseek-v4-pro | DeepSeek | 12.8 | percent |
One best compatible score per canonical product · higher is better
| Rank | Model | Weights | Evidence | Score |
|---|---|---|---|---|
| 01 | GPT-5.6 SolOpenAI | Closed weights | Public reference | 73% |
| 02 | Claude Fable 5Anthropic | Closed weights | Public reference | 70% |
| 03 | GPT-5.6 TerraOpenAI | Closed weights | Public reference | 69.6% |
| 04 | Claude Opus 5Anthropic | Closed weights | Public reference | 68.8% |
| 05 | Kimi K3Moonshot AI | Open weights | Source carrier | 67.5 USD |
| 06 | GPT-5.6 LunaOpenAI | Closed weights | Source carrier | 67.2 USD |
| 07 | GLM-5.3Z.AI | Open announced | Public reference | 66.9% |
| 08 | Grok 4.6xAI | Closed weights | Public reference | 65.9% |
| 09 | Gemini 3.7 FlashGoogle | Closed weights | Public reference | 65.3% |
| 10 | DeepSeek V4 Pro 0813DeepSeek | Open weights | Public reference | 62.7% |
Each distinct published measurement is retained. Same-snapshot canonical and configuration projections appear once.
| Published model / configuration | Score | Evidence & protocol | Source & dates |
|---|---|---|---|
DeepSeek-V4-Flash-0731 (provider-published configuration) as published in Qwen's Qwen3.8-Flash-Next comparison tableDeepSeek V4 Flash 0731Exact identitydeepseek-v4-flash-0731-deepseek-0813-release-unspecified Canonical product: deepseek-v4-flash-0731 | 54.4% | Source carrierprovider-reportedVersion & system1.1 mini-swe-agent | Qwen3.8-Flash-Next launch and official provider evaluationsObserved Checked |
Qwen3.8-27B (xhigh default thinking configuration) as published in Qwen's Qwen3.8-Flash-Next comparison tableQwen3.8-27BExact identityqwen-3-8-27b-xhigh Canonical product: qwen-3-8-27b | 42.2% | Relative comparisonprovider-reportedVersion & system1.1 mini-swe-agent | Qwen3.8-Flash-Next launch and official provider evaluationsObserved Checked |
Qwen3.7-Plus (provider-published configuration) as published in Qwen's Qwen3.8-Flash-Next comparison tableQwen3.7-PlusExact identityqwen-3-7-plus-unspecified Canonical product: qwen-3-7-plus | 16.5% | Relative comparisonprovider-reportedVersion & system1.1 mini-swe-agent | Qwen3.8-Flash-Next launch and official provider evaluationsObserved Checked |
DeepSeek Harness Minimal Mode; max effort; top_p=0.95; temperature=1.0DeepSeek V4 Flash Vision ExpExact identitydeepseek-v4-flash-vision-exp-max-harness Canonical product: deepseek-v4-flash-vision-exp | 59.3% | Public referenceprovider-reportedVersion & systemv1.1 DeepSeek Harness Minimal Mode | DeepSeek-V4-Flash-Vision-Exp release and provider evaluationObserved Checked |
Qwen3.8-27B (xhigh)Qwen3.8-27BExact identityqwen-3-8-27b-xhigh Canonical product: qwen-3-8-27b | 42.2% | Public referenceprovider-reportedVersion & system1.1 Qwen3.8-27B official text table | Qwen3.8-27B official model cardObserved Checked |
Qwen3.7-Plus as published by QwenQwen3.7-PlusExact identityqwen-3-7-plus-unspecified Canonical product: qwen-3-7-plus | 14.2% | Relative comparisonprovider-reportedVersion & system1.1 Qwen3.8-27B official text table | Qwen3.8-27B official model cardObserved Checked |
Qwen3.6-27B as published by QwenQwen3.6 27BExact identityqwen3-6-27b-default Canonical product: qwen3-6-27b | 13.3% | Relative comparisonprovider-reportedVersion & system1.1 Qwen3.8-27B official text table | Qwen3.8-27B official model cardObserved Checked |
GPT-5.6 Sol as published by Z.AIGPT-5.6 SolExact identitygpt-5-6-sol-max Canonical product: gpt-5-6-sol | 72.7% | Relative comparisonprovider-reportedVersion & system1.1 Z.AI GLM-5.3 official comparison table | GLM-5.3: Frontier Coding with Emergent Cyber CapabilitiesObserved Checked |
Claude Fable 5 w/ fallback as published by Z.AIClaude Fable 5Exact identityclaude-fable-5-max Canonical product: claude-fable-5 | 69.7% | Relative comparisonprovider-reportedVersion & system1.1 Z.AI GLM-5.3 official comparison table | GLM-5.3: Frontier Coding with Emergent Cyber CapabilitiesObserved Checked |
Kimi K3 as published by Z.AIKimi K3Exact identitykimi-k3-max Canonical product: kimi-k3 | 67.5% | Relative comparisonprovider-reportedVersion & system1.1 Z.AI GLM-5.3 official comparison table | GLM-5.3: Frontier Coding with Emergent Cyber CapabilitiesObserved Checked |
| Model | Result | Execution & attribution | Original source |
|---|---|---|---|
| GPT-5.6 SolDeepSWE 1.1 · Cognition release table | 72.7percent (published) | Release-table configuration; per-model effort and harness vary. See original footnotes.provider measured competitorSource-native comparison only. Different settings are not silently substituted into retained index protocols. | Cognition ↗Published 2026-09-10 · checked 2026-09-16 |
| SWE-1.7DeepSWE 1.1 · Cognition release table | 37.7percent (published) | Release-table configuration; per-model effort and harness vary. See original footnotes.provider subjectSource-native comparison only. Different settings are not silently substituted into retained index protocols. | Cognition ↗Published 2026-09-10 · checked 2026-09-16 |
| Kimi K3DeepSWE 1.1 · Cognition release table | 68.5percent (published) | Release-table configuration; per-model effort and harness vary. See original footnotes.provider measured competitorSource-native comparison only. Different settings are not silently substituted into retained index protocols. | Cognition ↗Published 2026-09-10 · checked 2026-09-16 |
| Grok 4.6DeepSWE 1.1 · Cognition release table | 67.5percent (published) | Release-table configuration; per-model effort and harness vary. See original footnotes.provider measured competitorSource-native comparison only. Different settings are not silently substituted into retained index protocols. | Cognition ↗Published 2026-09-10 · checked 2026-09-16 |
| SWE-2DeepSWE 1.1 · Cognition release table | 73percent (published) | Release-table configuration; per-model effort and harness vary. See original footnotes.provider subjectSource-native comparison only. Different settings are not silently substituted into retained index protocols. | Cognition ↗Published 2026-09-10 · checked 2026-09-16 |
| GPT-6 AstraDeepSWE 1.1 · Cognition release table | 74.1percent (published) | Release-table configuration; per-model effort and harness vary. See original footnotes.provider measured competitorSource-native comparison only. Different settings are not silently substituted into retained index protocols. | Cognition ↗Published 2026-09-10 · checked 2026-09-16 |
| Fable 5.1DeepSWE 1.1 · Cognition release table | 67.4percent (published) | Release-table configuration; per-model effort and harness vary. See original footnotes.provider measured competitorSource-native comparison only. Different settings are not silently substituted into retained index protocols. | Cognition ↗Published 2026-09-10 · checked 2026-09-16 |
From result to context
Contamination-resistant software engineering repair tasks across real repositories under a shared agent harness.
The summary shows one best compatible source result per canonical product. Versions, effort settings and execution systems remain attached to the underlying records.
Scores from different versions or harnesses may not be interchangeable. The published source rows preserve those distinctions and their original units. A source result is not automatically an input to a current index.
A reviewed track from this benchmark is represented in the following current indices. Each index applies its own protocol and eligibility rules.
Contamination risk: Low. Lifecycle: Active.