A task with its own rules.
A continuously updated benchmark for code generation, execution, self-repair and related tasks.
- Organisation
- LiveCodeBench team
- Version
- Current / rolling
Coding / Benchmark profile
A continuously updated benchmark for code generation, execution, self-repair and related tasks.
Observed results
20 ranked models · higher is better · labels show rank and score
| Model | Rank | Model ID | Provider | Score | Unit |
|---|---|---|---|---|---|
| Solar Open 2 250B | #1 | solar-open2-250b | Upstage | 92.4 | percent |
| Qwen3.8-Flash-Next | #2 | qwen-3-8-flash-next | Alibaba Cloud | 91.9 | percent |
| Qwen3.7-Max | #3 | qwen-3-7-max | Alibaba Cloud | 91.6 | percent |
| Qwen3.8-27B | #4 | qwen-3-8-27b | Alibaba Cloud | 90.3 | percent |
| Qwen3.7-Plus | #5 | qwen-3-7-plus | Alibaba Cloud | 89.6 | percent |
| GLM-4.7 | #6 | glm-4-7 | Z.AI | 89.418 | percent |
| GPT-5.2 | #7 | gpt-5-2 | OpenAI | 88.8889 | percent |
| Qwen3.6 Plus | #8 | qwen3-6-plus | Alibaba Cloud | 87.1 | percent |
| DeepSeek V3.2 | #9 | deepseek-v3-2 | DeepSeek | 86.2434 | percent |
| o4-mini | #10 | o4-mini | OpenAI | 85.9259 | percent |
| Kimi K2 Thinking | #11 | kimi-k2-thinking | Moonshot AI | 85.291 | percent |
| Kimi K2.5 | #12 | kimi-k2-5 | Moonshot AI | 85 | percent |
| Claude Opus 4.5 | #13 | claude-opus-4-5 | Anthropic | 84.8 | percent |
| GPT-5 | #14 | gpt-5 | OpenAI | 84.5503 | percent |
| GPT-5 Codex | #15 | gpt-5-codex | OpenAI | 84.0212 | percent |
| Qwen3.6 27B | #16 | qwen3-6-27b | Alibaba Cloud | 83.9 | percent |
| Grok 4 Fast | #17 | grok-4-fast | xAI | 83.1746 | percent |
| MiniMax M2 | #18 | minimax-m2 | MiniMax | 82.6455 | percent |
| Grok 4 | #19 | grok-4 | xAI | 81.9048 | percent |
| MiniMax M2.1 | #20 | minimax-m2-1 | MiniMax | 80.9524 | percent |
| o3 | #21 | o3 | OpenAI | 80.8466 | percent |
| Qwen3.6-35B-A3B | #22 | qwen3-6-35b-a3b | Alibaba Cloud | 80.4 | percent |
| Gemini 2.5 Pro | #23 | gemini-2-5-pro | 80.1058 | percent | |
| DeepSeek V3.1 Terminus | #24 | deepseek-v3-1-terminus | DeepSeek | 79.7884 | percent |
| DeepSeek R1 0528 (May '25) | #25 | deepseek-r1-aa-2 | DeepSeek | 73.1 | percent |
| Nova 2 Pro | #26 | nova-2-pro | Amazon | 73.0159 | percent |
| Gemini 2.5 Pro Preview (May' 25) | #27 | gemini-2-5-pro-05-06 | 71.8 | percent | |
| Claude Sonnet 4.5 | #28 | claude-sonnet-4-5 | Anthropic | 71.4286 | percent |
| Gemini 2.5 Flash | #29 | gemini-2-5-flash | 71.3228 | percent | |
| Exaone 4.0 32B | #30 | exaone-4-0-32b | LG AI Research | 70 | percent |
| Grok 3 mini | #31 | grok-3-mini | xAI | 69.6296 | percent |
| GLM-4.6 | #32 | glm-4-6 | Z.AI | 69.5238 | percent |
| GPT-5 mini | #33 | gpt-5-mini | OpenAI | 69.2063 | percent |
| o1 | #34 | o1 | OpenAI | 67.9365 | percent |
| Grok Code Fast 1 | #35 | grok-code-fast-1 | xAI | 65.7143 | percent |
| Claude Sonnet 4 | #36 | claude-sonnet-4 | Anthropic | 65.5026 | percent |
| Claude Opus 4.1 | #37 | claude-opus-4-1 | Anthropic | 65.3545 | percent |
| Claude Opus 4 | #38 | claude-opus-4 | Anthropic | 63.5979 | percent |
| Kimi K2 0905 | #39 | kimi-k2-0905 | Moonshot AI | 60.9524 | percent |
| DeepSeek V4 Pro | #40 | deepseek-v4-pro | DeepSeek | 56.8 | percent |
| DeepSeek V4 Flash | #41 | deepseek-v4-flash | DeepSeek | 55.2 | percent |
| Claude Haiku 4.5 | #42 | claude-haiku-4-5 | Anthropic | 51.11 | percent |
| Claude Sonnet 3.7 | #43 | claude-sonnet-3-7 | Anthropic | 47.3016 | percent |
| Mistral Large 3 | #44 | mistral-large-3 | Mistral AI | 46.46 | percent |
| DeepSeek V3 | #45 | deepseek-v3 | DeepSeek | 37.6 | percent |
| Claude 3.5 Sonnet (Oct '24) | #46 | claude-35-sonnet | Anthropic | 36.4 | percent |
| GPT-4o (Aug '24) | #47 | gpt-4o-2024-08-06 | OpenAI | 29.5 | percent |
| GPT-4 Turbo | #48 | gpt-4-turbo | OpenAI | 28.7 | percent |
| GPT-4o mini | #49 | gpt-4o-mini | OpenAI | 27.5 | percent |
| Claude 3 Haiku | #50 | claude-3-haiku | Anthropic | 20.2 | percent |
One best compatible score per canonical product · higher is better
| Rank | Model | Weights | Evidence | Score |
|---|---|---|---|---|
| 01 | Solar Open 2 250BUpstage | Open weights | Direct source result | 92.4% |
| 02 | Qwen3.8-Flash-NextAlibaba Cloud | Open weights | Direct source result | 91.9% |
| 03 | Qwen3.7-MaxAlibaba Cloud | Closed weights | Source carrier | 91.6% |
| 04 | Qwen3.8-27BAlibaba Cloud | Open weights | Direct source result | 90.3% |
| 05 | Qwen3.7-PlusAlibaba Cloud | Closed weights | Source carrier | 89.6% |
| 06 | GLM-4.7Z.AI | Open weights | Public reference | 89.4% |
| 07 | GPT-5.2OpenAI | Closed weights | Public reference | 88.9% |
| 08 | Qwen3.6 PlusAlibaba Cloud | Closed weights | Source carrier | 87.1% |
| 09 | DeepSeek V3.2DeepSeek | Open weights | Public reference | 86.2% |
| 10 | o4-miniOpenAI | Closed weights | Public reference | 85.9% |
Each distinct published measurement is retained. Same-snapshot canonical and configuration projections appear once.
| Published model / configuration | Score | Evidence & protocol | Source & dates |
|---|---|---|---|
Qwen3.8-Flash-Next (xhigh default thinking configuration)Qwen3.8-Flash-NextExact identityqwen-3-8-flash-next-xhigh Canonical product: qwen-3-8-flash-next | 91.9% | Direct source resultsource-checkedVersion & systemv6 system:qwen3-8-flash-next:lcb-v6 | Qwen3.8-Flash-Next current official launch pageObserved Checked |
DeepSeek-V4-Flash-0731 (provider-published configuration) as published in Qwen's Qwen3.8-Flash-Next comparison tableDeepSeek V4 Flash 0731Exact identitydeepseek-v4-flash-0731-deepseek-0813-release-unspecified Canonical product: deepseek-v4-flash-0731 | 90.6% | Relative comparisonprovider-reportedVersion & systemv6 system:qwen3-8-flash-next-comparison:cell:language:lcb-v6:3 | Qwen3.8-Flash-Next launch and official provider evaluationsObserved Checked |
Qwen3.8-27B (xhigh default thinking configuration) as published in Qwen's Qwen3.8-Flash-Next comparison tableQwen3.8-27BExact identityqwen-3-8-27b-xhigh Canonical product: qwen-3-8-27b | 90.3% | Relative comparisonprovider-reportedVersion & systemv6 system:qwen3-8-flash-next-comparison:cell:language:lcb-v6:1 | Qwen3.8-Flash-Next launch and official provider evaluationsObserved Checked |
Qwen3.7-Plus (provider-published configuration) as published in Qwen's Qwen3.8-Flash-Next comparison tableQwen3.7-PlusExact identityqwen-3-7-plus-unspecified Canonical product: qwen-3-7-plus | 89.6% | Relative comparisonprovider-reportedVersion & systemv6 system:qwen3-8-flash-next-comparison:cell:language:lcb-v6:2 | Qwen3.8-Flash-Next launch and official provider evaluationsObserved Checked |
Claude Opus 4.6 (Max) as published in Qwen's Qwen3.8-Flash-Next comparison tableClaude Opus 4.6Exact identityclaude-opus-4-6-max Canonical product: claude-opus-4-6 | 88.8% | Relative comparisonprovider-reportedVersion & systemv6 system:qwen3-8-flash-next-comparison:cell:language:lcb-v6:4 | Qwen3.8-Flash-Next launch and official provider evaluationsObserved Checked |
Solar Open 2 250B (high)Solar Open 2 250BExact identitysolar-open2-250b-high Canonical product: solar-open2-250b | 92.4% | Direct source resultprovider-reportedVersion & systemv6 Solar Open 2 card English table | Solar Open 2 250B model cardObserved Checked |
DeepSeek V4 Flash max as published by UpstageDeepSeek V4 FlashExact identitydeepseek-v4-flash-max Canonical product: deepseek-v4-flash | 92.3% | Relative comparisonprovider-reportedVersion & systemv6 Solar Open 2 card English table | Solar Open 2 250B model cardObserved Checked |
Qwen3.8-27B (xhigh)Qwen3.8-27BExact identityqwen-3-8-27b-xhigh Canonical product: qwen-3-8-27b | 90.3% | Direct source resultprovider-reportedVersion & systemv6 Qwen3.8-27B official text table | Qwen3.8-27B official model cardObserved Checked |
Qwen3.7-Plus as published by QwenQwen3.7-PlusExact identityqwen-3-7-plus-unspecified Canonical product: qwen-3-7-plus | 89.6% | Relative comparisonprovider-reportedVersion & systemv6 Qwen3.8-27B official text table | Qwen3.8-27B official model cardObserved Checked |
MiMo-V2.5 as published by UpstageMiMo-V2.5Exact identitymimo-v2-5-solar-open2-unspecified Canonical product: mimo-v2-5 | 89.1% | Relative comparisonprovider-reportedVersion & systemv6 Solar Open 2 card English table | Solar Open 2 250B model cardObserved Checked |
From result to context
A continuously updated benchmark for code generation, execution, self-repair and related tasks.
The summary shows one best compatible source result per canonical product. Versions, effort settings and execution systems remain attached to the underlying records.
Scores from different versions or harnesses may not be interchangeable. The published source rows preserve those distinctions and their original units. A source result is not automatically an input to a current index.
This catalogue entry preserves available source evidence. Current indices admit only their specifically reviewed tracks and configurations.
Contamination risk: Medium. Lifecycle: Rolling.