A task with its own rules.
A repository-understanding benchmark that measures whether models can map natural-language requests onto the right code locations and system changes.
- Organisation
- MiniMax
- Version
- 2026
Coding / Benchmark profile
A repository-understanding benchmark that measures whether models can map natural-language requests onto the right code locations and system changes.
Observed results
20 ranked models · higher is better · labels show rank and score
| Model | Rank | Model ID | Provider | Score | Unit |
|---|---|---|---|---|---|
| Claude Opus 4.8 | #1 | claude-opus-4-8 | Anthropic | 69.7 | percent |
| DeepSeek V4 Pro 0813 | #2 | deepseek-v4-pro-0813 | DeepSeek | 61.5 | percent |
| GLM-5.3 | #3 | glm-5-3 | Z.AI | 58 | percent |
| DeepSeek V4 Flash Vision Exp | #4 | deepseek-v4-flash-vision-exp | DeepSeek | 57.7 | percent |
| DeepSeek V4 Flash | #5 | deepseek-v4-flash | DeepSeek | 54.2 | percent |
| DeepSeek V4 Flash 0731 | #6 | deepseek-v4-flash-0731 | DeepSeek | 54.2 | percent |
| GLM-5.2 | #7 | glm-5-2 | Z.AI | 48.9 | percent |
| Ornith-1.0-397B | #8 | ornith-1-0-397b | DeepReinforce AI | 48.2 | percent |
| Qwen3.7-Max | #9 | qwen-3-7-max | Alibaba Cloud | 47.2 | percent |
| Claude Opus 4.5 | #10 | claude-opus-4-5 | Anthropic | 43.2 | percent |
| Qwen 3.6 Max (preview) | #11 | qwen3-6-max-preview | Alibaba Cloud | 42.9 | percent |
| GLM-5.1 | #12 | glm-5-1 | Z.AI | 42.7 | percent |
| Qwen3.8-27B | #13 | qwen-3-8-27b | Alibaba Cloud | 42.3 | percent |
| MiniMax M3 | #14 | minimax-m3 | MiniMax | 42.13 | percent |
| Qwen3.7-Plus | #15 | qwen-3-7-plus | Alibaba Cloud | 41.1 | percent |
| MiniMax M2.7 | #16 | minimax-m2-7 | MiniMax | 39.8 | percent |
| DeepSeek V4 Pro | #17 | deepseek-v4-pro | DeepSeek | 38.5 | percent |
| Qwen3.6 27B | #18 | qwen3-6-27b | Alibaba Cloud | 36.2 | percent |
| Ornith-1.0-35B | #19 | ornith-1-0-35b | DeepReinforce AI | 34.6 | percent |
| Qwen3.6-35B-A3B | #20 | qwen3-6-35b-a3b | Alibaba Cloud | 29.4 | percent |
| Ornith-1.0-9B | #21 | ornith-1-0-9b | DeepReinforce AI | 27.2 | percent |
One best compatible score per canonical product · higher is better
| Rank | Model | Weights | Evidence | Score |
|---|---|---|---|---|
| 01 | Claude Opus 4.8Anthropic | Closed weights | Public reference | 69.7% |
| 02 | DeepSeek V4 Pro 0813DeepSeek | Open weights | Public reference | 61.5% |
| 03 | GLM-5.3Z.AI | Open announced | Public reference | 58% |
| 04 | DeepSeek V4 Flash Vision ExpDeepSeek | Not documented | Public reference | 57.7% |
| 05 | DeepSeek V4 FlashDeepSeek | Open weights | Source carrier | 54.2% |
| 06 | DeepSeek V4 Flash 0731DeepSeek | Open weights | Public reference | 54.2% |
| 07 | GLM-5.2Z.AI | Open weights | Public reference | 48.9% |
| 08 | Ornith-1.0-397BDeepReinforce AI | Open weights | Source carrier | 48.2% |
| 09 | Qwen3.7-MaxAlibaba Cloud | Closed weights | Source carrier | 47.2% |
| 10 | Claude Opus 4.5Anthropic | Closed weights | Source carrier | 43.2% |
Each distinct published measurement is retained. Same-snapshot canonical and configuration projections appear once.
| Published model / configuration | Score | Evidence & protocol | Source & dates |
|---|---|---|---|
DeepSeek-V4-Flash-0731 (provider-published configuration) as published in Qwen's Qwen3.8-Flash-Next comparison tableDeepSeek V4 Flash 0731Exact identitydeepseek-v4-flash-0731-deepseek-0813-release-unspecified Canonical product: deepseek-v4-flash-0731 | 54.2% | Public referenceprovider-reportedVersion & system2026 Claude Code | Qwen3.8-Flash-Next launch and official provider evaluationsObserved Checked |
Claude Opus 4.6 (Max) as published in Qwen's Qwen3.8-Flash-Next comparison tableClaude Opus 4.6Exact identityclaude-opus-4-6-max Canonical product: claude-opus-4-6 | 47.6% | Public referenceprovider-reportedVersion & system2026 Claude Code | Qwen3.8-Flash-Next launch and official provider evaluationsObserved Checked |
Qwen3.8-27B (xhigh default thinking configuration) as published in Qwen's Qwen3.8-Flash-Next comparison tableQwen3.8-27BExact identityqwen-3-8-27b-xhigh Canonical product: qwen-3-8-27b | 42.3% | Public referenceprovider-reportedVersion & system2026 Claude Code | Qwen3.8-Flash-Next launch and official provider evaluationsObserved Checked |
Qwen3.7-Plus (provider-published configuration) as published in Qwen's Qwen3.8-Flash-Next comparison tableQwen3.7-PlusExact identityqwen-3-7-plus-unspecified Canonical product: qwen-3-7-plus | 41.1% | Public referenceprovider-reportedVersion & system2026 Claude Code | Qwen3.8-Flash-Next launch and official provider evaluationsObserved Checked |
DeepSeek Harness Minimal Mode; max effort; top_p=0.95; temperature=1.0DeepSeek V4 Flash Vision ExpExact identitydeepseek-v4-flash-vision-exp-max-harness Canonical product: deepseek-v4-flash-vision-exp | 57.7% | Public referenceprovider-reportedVersion & system2026 DeepSeek Harness Minimal Mode | DeepSeek-V4-Flash-Vision-Exp release and provider evaluationObserved Checked |
Claude Opus 4.6 Max as published by QwenClaude Opus 4.6Exact identityclaude-opus-4-6-max Canonical product: claude-opus-4-6 | 47.6% | Public referenceprovider-reportedVersion & system2026 Qwen3.8-27B official text table | Qwen3.8-27B official model cardObserved Checked |
Qwen3.8-27B (xhigh)Qwen3.8-27BExact identityqwen-3-8-27b-xhigh Canonical product: qwen-3-8-27b | 42.3% | Public referenceprovider-reportedVersion & system2026 Qwen3.8-27B official text table | Qwen3.8-27B official model cardObserved Checked |
Qwen3.7-Plus as published by QwenQwen3.7-PlusExact identityqwen-3-7-plus-unspecified Canonical product: qwen-3-7-plus | 41.1% | Public referenceprovider-reportedVersion & system2026 Qwen3.8-27B official text table | Qwen3.8-27B official model cardObserved Checked |
Qwen3.6-27B as published by QwenQwen3.6 27BExact identityqwen3-6-27b-default Canonical product: qwen3-6-27b | 36.2% | Public referenceprovider-reportedVersion & system2026 Qwen3.8-27B official text table | Qwen3.8-27B official model cardObserved Checked |
Claude Opus 4.8 as published by Z.AIClaude Opus 4.8Exact identityclaude-opus-4-8-max Canonical product: claude-opus-4-8 | 69.7% | Public referenceprovider-reportedVersion & system2026 Z.AI GLM-5.3 official comparison table | GLM-5.3: Frontier Coding with Emergent Cyber CapabilitiesObserved Checked |
From result to context
A repository-understanding benchmark that measures whether models can map natural-language requests onto the right code locations and system changes.
The summary shows one best compatible source result per canonical product. Versions, effort settings and execution systems remain attached to the underlying records.
Scores from different versions or harnesses may not be interchangeable. The published source rows preserve those distinctions and their original units. A source result is not automatically an input to a current index.
This catalogue entry preserves available source evidence. Current indices admit only their specifically reviewed tracks and configurations.
Contamination risk: Unknown. Lifecycle: Active.