A task with its own rules.
A human-validated subset of real GitHub software-engineering issue resolution tasks.
- Organisation
- SWE-bench authors
- Version
- Verified
Coding / Benchmark profile
A human-validated subset of real GitHub software-engineering issue resolution tasks.
Observed results
20 ranked models · higher is better · labels show rank and score
| Model | Rank | Model ID | Provider | Score | Unit |
|---|---|---|---|---|---|
| Claude Opus 5 | #1 | claude-opus-5 | Anthropic | 96 | percent |
| Claude Mythos 5 | #2 | claude-mythos-5 | Anthropic | 95.5 | percent |
| Claude Fable 5 | #3 | claude-fable-5 | Anthropic | 95 | percent |
| Claude Opus 4.8 | #4 | claude-opus-4-8 | Anthropic | 88.6 | percent |
| Claude Opus 4.7 | #5 | claude-opus-4-7 | Anthropic | 87.6 | percent |
| Claude Sonnet 5 | #6 | claude-sonnet-5 | Anthropic | 85.2 | percent |
| GPT-5.3-Codex | #7 | gpt-5-3-codex | OpenAI | 85 | percent |
| Ornith-1.0-397B | #8 | ornith-1-0-397b | DeepReinforce AI | 82.4 | percent |
| Claude Opus 4.5 | #9 | claude-opus-4-5 | Anthropic | 80.9 | percent |
| Claude Opus 4.6 | #10 | claude-opus-4-6 | Anthropic | 80.84 | percent |
| DeepSeek V4 Pro | #11 | deepseek-v4-pro | DeepSeek | 80.6 | percent |
| MiniMax M3 | #12 | minimax-m3 | MiniMax | 80.5 | percent |
| Qwen3.7-Max | #13 | qwen-3-7-max | Alibaba Cloud | 80.4 | percent |
| Inkling-Small | #14 | inkling-small | Thinking Machines Lab | 80.2 | percent |
| Kimi K2.6 | #15 | kimi-k2-6 | Moonshot AI | 80.2 | percent |
| GPT-5.2 | #16 | gpt-5-2 | OpenAI | 80 | percent |
| Claude Sonnet 4.6 | #17 | claude-sonnet-4-6 | Anthropic | 79.6 | percent |
| DeepSeek V4 Flash | #18 | deepseek-v4-flash | DeepSeek | 79 | percent |
| Qwen3.6 Plus | #19 | qwen3-6-plus | Alibaba Cloud | 78.8 | percent |
| Hy3 | #20 | hy3 | Tencent | 78 | percent |
| MiMo-V2-Pro | #21 | mimo-v2-pro | Xiaomi | 78 | percent |
| GLM-5 | #22 | glm-5 | Z.AI | 77.8 | percent |
| Qwen3.7-Plus | #23 | qwen-3-7-plus | Alibaba Cloud | 77.7 | percent |
| Inkling | #24 | inkling | Thinking Machines Lab | 77.6 | percent |
| Mistral Medium 3.5 128B | #25 | mistral-medium-3-5-128b | Mistral AI | 77.6 | percent |
| Muse Spark | #26 | muse-spark | Meta | 77.4 | percent |
| Claude Sonnet 4.5 | #27 | claude-sonnet-4-5 | Anthropic | 77.2 | percent |
| Qwen3.6 27B | #28 | qwen3-6-27b | Alibaba Cloud | 77.2 | percent |
| Kimi K2.5 | #29 | kimi-k2-5 | Moonshot AI | 76.8 | percent |
| Grok 4.20 | #30 | grok-4-20 | xAI | 76.7 | percent |
| Qwen3.5 397B A17B | #31 | qwen3-5-397b | Alibaba Cloud | 76.2 | percent |
| Muse Glimmer 30B | #32 | muse-glimmer-30b | Meta | 76 | percent |
| Ornith-1.0-35B | #33 | ornith-1-0-35b | DeepReinforce AI | 75.6 | percent |
| MiMo-V2-Omni | #34 | mimo-v2-omni | Xiaomi | 74.8 | percent |
| Laguna M.1 | #35 | laguna-m-1 | Poolside | 74.6 | percent |
| Claude 4.1 Opus | #36 | claude-4-1-opus | Anthropic | 74.5 | percent |
| Hy3 Preview | #37 | hy3-preview | Tencent | 74.4 | percent |
| GLM-4.7 | #38 | glm-4-7 | Z.AI | 73.8 | percent |
| MAI-Thinking-1 | #39 | mai-thinking-1 | Microsoft | 73.5 | percent |
| Qwen3.6-35B-A3B | #40 | qwen3-6-35b-a3b | Alibaba Cloud | 73.4 | percent |
| Claude Haiku 4.5 | #41 | claude-haiku-4-5 | Anthropic | 73.3 | percent |
| Claude 4 Sonnet | #42 | claude-4-sonnet | Anthropic | 72.7 | percent |
| Qwen3.5 27B | #43 | qwen3-5-27b | Alibaba Cloud | 72.4 | percent |
| Qwen3.5-122B-A10B | #44 | qwen3-5-122b-a10b | Alibaba Cloud | 72 | percent |
| Nemotron 3 Ultra | #45 | nemotron-3-ultra | NVIDIA | 71.9 | percent |
| Laguna XS 2.1 | #46 | laguna-xs-2-1 | Poolside | 70.9 | percent |
| Grok Code Fast 1 | #47 | grok-code-fast-1 | xAI | 70.8 | percent |
| Solar Open 2 250B | #48 | solar-open2-250b | Upstage | 70.4 | percent |
| Laguna XS.2 | #49 | laguna-xs-2 | Poolside | 69.9 | percent |
| Ornith-1.0-9B | #50 | ornith-1-0-9b | DeepReinforce AI | 69.4 | percent |
| Qwen3.5-35B-A3B | #51 | qwen3-5-35b-a3b | Alibaba Cloud | 69.2 | percent |
| Gemini 2.5 Pro | #52 | gemini-2-5-pro | 63.8 | percent | |
| GPT-4.1 | #53 | gpt-4-1 | OpenAI | 54.6 | percent |
| ZAYA1-74B-Preview | #54 | zaya1-74b-preview | Zyphra | 53.2 | percent |
| NVIDIA Nemotron 3.5 Lightning 30B-A3B | #55 | nemotron-3-5-lightning-30b-a3b | NVIDIA | 51.56 | percent |
| o3-mini | #56 | o3-mini | OpenAI | 49.3 | percent |
| Claude 3.5 Sonnet | #57 | claude-3-5-sonnet | Anthropic | 49 | percent |
| DeepSeek V3 | #58 | deepseek-v3 | DeepSeek | 42 | percent |
| GPT-4.1 mini | #59 | gpt-4-1-mini | OpenAI | 23.6 | percent |
One best compatible score per canonical product · higher is better
| Rank | Model | Weights | Evidence | Score |
|---|---|---|---|---|
| 01 | Claude Opus 5Anthropic | Closed weights | Source carrier | 96% |
| 02 | Claude Mythos 5Anthropic | Closed weights | Source carrier | 95.5% |
| 03 | Claude Fable 5Anthropic | Closed weights | Source carrier | 95% |
| 04 | Claude Opus 4.8Anthropic | Closed weights | Source carrier | 88.6% |
| 05 | Claude Opus 4.7Anthropic | Closed weights | Source carrier | 87.6% |
| 06 | Claude Sonnet 5Anthropic | Closed weights | Source carrier | 85.2% |
| 07 | GPT-5.3-CodexOpenAI | Closed weights | Source carrier | 85% |
| 08 | Ornith-1.0-397BDeepReinforce AI | Open weights | Source carrier | 82.4% |
| 09 | Claude Opus 4.5Anthropic | Closed weights | Source carrier | 80.9% |
| 10 | Claude Opus 4.6Anthropic | Closed weights | Source carrier | 80.8% |
Each distinct published measurement is retained. Same-snapshot canonical and configuration projections appear once.
| Published model / configuration | Score | Evidence & protocol | Source & dates |
|---|---|---|---|
DeepSeek V4 Flash max as published by UpstageDeepSeek V4 FlashExact identitydeepseek-v4-flash-max Canonical product: deepseek-v4-flash | 73.8% | Relative comparisonprovider-reportedVersion & systemVerified Solar Open 2 card English table | Solar Open 2 250B model cardObserved Checked |
MiMo-V2.5 as published by UpstageMiMo-V2.5Exact identitymimo-v2-5-solar-open2-unspecified Canonical product: mimo-v2-5 | 73% | Relative comparisonprovider-reportedVersion & systemVerified Solar Open 2 card English table | Solar Open 2 250B model cardObserved Checked |
Solar Open 2 250B (high)Solar Open 2 250BExact identitysolar-open2-250b-high Canonical product: solar-open2-250b | 70.4% | Public referenceprovider-reportedVersion & systemVerified Solar Open 2 card English table | Solar Open 2 250B model cardObserved Checked |
Mistral Medium 3.5 as published by UpstageMistral Medium 3.5Exact identitymistral-medium-3-5-high Canonical product: mistral-medium-3-5 | 69.6% | Relative comparisonprovider-reportedVersion & systemVerified Solar Open 2 card English table | Solar Open 2 250B model cardObserved Checked |
Solar Open 100B as published by UpstageSolar Open 100B (Reasoning)Exact identitysolar-open-100b-reasoning-high Canonical product: solar-open-100b-reasoning | 15.4% | Relative comparisonprovider-reportedVersion & systemVerified Solar Open 2 card English table | Solar Open 2 250B model cardObserved Checked |
Command A+ as published by UpstageCommand A+Exact identitycommand-a-plus-solar-open2-unspecified Canonical product: command-a-plus | 14.4% | Relative comparisonprovider-reportedVersion & systemVerified Solar Open 2 card English table | Solar Open 2 250B model cardObserved Checked |
NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16; NVIDIA release evaluation; temperature=1.0; top_p=0.95NVIDIA Nemotron 3.5 Lightning 30B-A3BExact identitynemotron-3-5-lightning-30b-a3b-default Canonical product: nemotron-3-5-lightning-30b-a3b | 51.6% | Public referenceprovider-reportedVersion & systemVerified NVIDIA NeMo Gym / NeMo Evaluator SDK consistent release harness | NVIDIA Nemotron 3.5 Lightning 30B-A3B BF16 model cardObserved Checked |
Muse Glimmer-30B; high reasoning; temperature=1.0; top_p=0.95; top_k=64Muse Glimmer 30BExact identitymuse-glimmer-30b-high Canonical product: muse-glimmer-30b | 76% | Public referenceprovider-reportedVersion & systemVerified / 500 tasks Meta agentic coding scaffold, averaged across four runs | Muse Glimmer Evaluation MethodologyObserved Checked |
Exact BenchLM registry variant Claude Opus 5; bulk export does not retain a complete upstream harness configuration.Claude Opus 5Exact identitySource label without a registered configuration ID Canonical product: claude-opus-5 | 96% | Source carriersource-checkedVersion & system2024 Source-native system | BenchLM public datasets — 2026-08-01Observed Checked |
Exact BenchLM registry variant Claude Mythos 5; bulk export does not retain a complete upstream harness configuration.Claude Mythos 5Exact identitySource label without a registered configuration ID Canonical product: claude-mythos-5 | 95.5% | Source carriersource-checkedVersion & system2024 Source-native system | BenchLM public datasets — 2026-08-01Observed Checked |
From result to context
A human-validated subset of real GitHub software-engineering issue resolution tasks.
The summary shows one best compatible source result per canonical product. Versions, effort settings and execution systems remain attached to the underlying records.
Scores from different versions or harnesses may not be interchangeable. The published source rows preserve those distinctions and their original units. A source result is not automatically an input to a current index.
This catalogue entry preserves available source evidence. Current indices admit only their specifically reviewed tracks and configurations.
Contamination risk: Medium. Lifecycle: Active.