A task with its own rules.
A multilingual extension of SWE-bench covering 300 problems across 9 programming languages, testing code generation and bug fixing beyond Python.
- Organisation
- SWE-bench team
- Version
- 2025
Knowledge / Benchmark profile
A multilingual extension of SWE-bench covering 300 problems across 9 programming languages, testing code generation and bug fixing beyond Python.
Observed results
20 ranked models · higher is better · labels show rank and score
| Model | Rank | Model ID | Provider | Score | Unit |
|---|---|---|---|---|---|
| Claude Mythos 5 | #1 | claude-mythos-5 | Anthropic | 92.2 | percent |
| Claude Opus 5 | #2 | claude-opus-5 | Anthropic | 89.5 | percent |
| Claude Opus 4.8 | #3 | claude-opus-4-8 | Anthropic | 84.4 | percent |
| Composer 2.5 | #4 | composer-2-5 | Cursor | 79.8 | percent |
| Ornith-1.0-397B | #5 | ornith-1-0-397b | DeepReinforce AI | 78.9 | percent |
| Laguna S 2.1 | #6 | laguna-s-2-1 | Poolside | 78.5 | percent |
| Claude Sonnet 5 | #7 | claude-sonnet-5 | Anthropic | 78.3 | percent |
| Qwen3.7-Max | #8 | qwen-3-7-max | Alibaba Cloud | 78.3 | percent |
| Grok 4.5 | #9 | grok-4-5 | xAI | 78 | percent |
| SWE-1.7 | #10 | swe-1-7 | Cognition | 77.8 | percent |
| Claude Opus 4.5 | #11 | claude-opus-4-5 | Anthropic | 77.5 | percent |
| LongCat-2.0 | #12 | longcat-2-0 | Meituan | 77.3 | percent |
| Kimi K2.6 | #13 | kimi-k2-6 | Moonshot AI | 76.7 | percent |
| MiniMax M2.7 | #14 | minimax-m2-7 | MiniMax | 76.5 | percent |
| DeepSeek V4 Pro | #15 | deepseek-v4-pro | DeepSeek | 76.2 | percent |
| Qwen3.7-Plus | #16 | qwen-3-7-plus | Alibaba Cloud | 75.8 | percent |
| Qwen3.6 Plus | #17 | qwen3-6-plus | Alibaba Cloud | 73.8 | percent |
| Composer 2 | #18 | composer-2 | Cursor | 73.7 | percent |
| DeepSeek V4 Flash | #19 | deepseek-v4-flash | DeepSeek | 73.3 | percent |
| GLM-5 | #20 | glm-5 | Z.AI | 73.3 | percent |
| Kimi K2.5 | #21 | kimi-k2-5 | Moonshot AI | 73 | percent |
| Qwen3.6 27B | #22 | qwen3-6-27b | Alibaba Cloud | 71.3 | percent |
| Ornith-1.0-35B | #23 | ornith-1-0-35b | DeepReinforce AI | 69.3 | percent |
| Nemotron 3 Ultra | #24 | nemotron-3-ultra | NVIDIA | 67.7 | percent |
| Qwen3.6-35B-A3B | #25 | qwen3-6-35b-a3b | Alibaba Cloud | 67.2 | percent |
| Laguna M.1 | #26 | laguna-m-1 | Poolside | 63.1 | percent |
| Laguna XS 2.1 | #27 | laguna-xs-2-1 | Poolside | 63.1 | percent |
| Laguna XS.2 | #28 | laguna-xs-2 | Poolside | 57.7 | percent |
| Ornith-1.0-9B | #29 | ornith-1-0-9b | DeepReinforce AI | 52 | percent |
| NVIDIA Nemotron 3.5 Lightning 30B-A3B | #30 | nemotron-3-5-lightning-30b-a3b | NVIDIA | 39.33 | percent |
One best compatible score per canonical product · higher is better
| Rank | Model | Weights | Evidence | Score |
|---|---|---|---|---|
| 01 | Claude Mythos 5Anthropic | Closed weights | Source carrier | 92.2% |
| 02 | Claude Opus 5Anthropic | Closed weights | Source carrier | 89.5% |
| 03 | Claude Opus 4.8Anthropic | Closed weights | Source carrier | 84.4% |
| 04 | Composer 2.5Cursor | Closed weights | Source carrier | 79.8% |
| 05 | Ornith-1.0-397BDeepReinforce AI | Open weights | Source carrier | 78.9% |
| 06 | Laguna S 2.1Poolside | Open weights | Source carrier | 78.5% |
| 07 | Claude Sonnet 5Anthropic | Closed weights | Source carrier | 78.3% |
| 08 | Qwen3.7-MaxAlibaba Cloud | Closed weights | Source carrier | 78.3% |
| 09 | Grok 4.5xAI | Closed weights | Source carrier | 78% |
| 10 | SWE-1.7Cognition | Closed weights | Source carrier | 77.8% |
Each distinct published measurement is retained. Same-snapshot canonical and configuration projections appear once.
| Published model / configuration | Score | Evidence & protocol | Source & dates |
|---|---|---|---|
Claude Opus 4.6 (Max) as published in Qwen's Qwen3.8-Flash-Next comparison tableClaude Opus 4.6Exact identityclaude-opus-4-6-max Canonical product: claude-opus-4-6 | 77.5% | Public referenceprovider-reportedVersion & system2025 mini-swe-agent | Qwen3.8-Flash-Next launch and official provider evaluationsObserved Checked |
Qwen3.7-Plus (provider-published configuration) as published in Qwen's Qwen3.8-Flash-Next comparison tableQwen3.7-PlusExact identityqwen-3-7-plus-unspecified Canonical product: qwen-3-7-plus | 75.8% | Public referenceprovider-reportedVersion & system2025 mini-swe-agent | Qwen3.8-Flash-Next launch and official provider evaluationsObserved Checked |
Qwen3.8-27B (xhigh default thinking configuration) as published in Qwen's Qwen3.8-Flash-Next comparison tableQwen3.8-27BExact identityqwen-3-8-27b-xhigh Canonical product: qwen-3-8-27b | 73.8% | Public referenceprovider-reportedVersion & system2025 mini-swe-agent | Qwen3.8-Flash-Next launch and official provider evaluationsObserved Checked |
NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16; NVIDIA release evaluation; temperature=1.0; top_p=0.95NVIDIA Nemotron 3.5 Lightning 30B-A3BExact identitynemotron-3-5-lightning-30b-a3b-default Canonical product: nemotron-3-5-lightning-30b-a3b | 39.3% | Public referenceprovider-reportedVersion & system2025 NVIDIA NeMo Gym / NeMo Evaluator SDK consistent release harness | NVIDIA Nemotron 3.5 Lightning 30B-A3B BF16 model cardObserved Checked |
Exact BenchLM registry variant Claude Mythos 5; bulk export does not retain a complete upstream harness configuration.Claude Mythos 5Exact identitySource label without a registered configuration ID Canonical product: claude-mythos-5 | 92.2% | Source carriersource-checkedVersion & system2025 Source-native system | BenchLM public datasets — 2026-08-01Observed Checked |
Exact BenchLM registry variant Claude Opus 5; bulk export does not retain a complete upstream harness configuration.Claude Opus 5Exact identitySource label without a registered configuration ID Canonical product: claude-opus-5 | 89.5% | Source carriersource-checkedVersion & system2025 Source-native system | BenchLM public datasets — 2026-08-01Observed Checked |
Exact BenchLM registry variant Claude Opus 4.8; bulk export does not retain a complete upstream harness configuration.Claude Opus 4.8Exact identitySource label without a registered configuration ID Canonical product: claude-opus-4-8 | 84.4% | Source carriersource-checkedVersion & system2025 Source-native system | BenchLM public datasets — 2026-08-01Observed Checked |
Exact BenchLM registry variant Composer 2.5; bulk export does not retain a complete upstream harness configuration.Composer 2.5Exact identitySource label without a registered configuration ID Canonical product: composer-2-5 | 79.8% | Source carriersource-checkedVersion & system2025 Source-native system | BenchLM public datasets — 2026-08-01Observed Checked |
Exact BenchLM registry variant Ornith-1.0-397B; bulk export does not retain a complete upstream harness configuration.Ornith-1.0-397BExact identitySource label without a registered configuration ID Canonical product: ornith-1-0-397b | 78.9% | Source carriersource-checkedVersion & system2025 Source-native system | BenchLM public datasets — 2026-08-01Observed Checked |
Exact BenchLM registry variant Laguna S 2.1; bulk export does not retain a complete upstream harness configuration.Laguna S 2.1Exact identitySource label without a registered configuration ID Canonical product: laguna-s-2-1 | 78.5% | Source carriersource-checkedVersion & system2025 Source-native system | BenchLM public datasets — 2026-08-01Observed Checked |
From result to context
A multilingual extension of SWE-bench covering 300 problems across 9 programming languages, testing code generation and bug fixing beyond Python.
The summary shows one best compatible source result per canonical product. Versions, effort settings and execution systems remain attached to the underlying records.
Scores from different versions or harnesses may not be interchangeable. The published source rows preserve those distinctions and their original units. A source result is not automatically an input to a current index.
This catalogue entry preserves available source evidence. Current indices admit only their specifically reviewed tracks and configurations.
Contamination risk: Unknown. Lifecycle: Active.