A task with its own rules.
A frequently refreshed benchmark suite designed to limit contamination and cover multiple capabilities.
- Organisation
- LiveBench team
- Version
- Current / rolling
Reasoning / Benchmark profile
A frequently refreshed benchmark suite designed to limit contamination and cover multiple capabilities.
Observed results
20 ranked models · higher is better · labels show rank and score
| Model | Rank | Model ID | Provider | Score | Unit |
|---|---|---|---|---|---|
| GPT-5.6 Sol | #1 | gpt-5-6-sol | OpenAI | 91.65375 | percent |
| Claude Opus 5 | #2 | claude-opus-5 | Anthropic | 91.2115 | percent |
| Kimi K3 | #3 | kimi-k3 | Moonshot AI | 90.673 | percent |
| GPT-5.6 Terra | #4 | gpt-5-6-terra | OpenAI | 90.6345 | percent |
| Grok 4.6 | #5 | grok-4-6 | xAI | 90.5095 | percent |
| Muse Spark 1.2 | #6 | muse-spark-1-2 | Meta | 90.00475 | percent |
| Claude Fable 5 | #7 | claude-fable-5 | Anthropic | 89.65375 | percent |
| GPT-5.5 | #8 | gpt-5-5 | OpenAI | 89.65375 | percent |
| Claude Opus 4.8 | #9 | claude-opus-4-8 | Anthropic | 89.19225 | percent |
| Claude Sonnet 5 | #10 | claude-sonnet-5 | Anthropic | 88.69225 | percent |
| Claude Opus 4.6 | #11 | claude-opus-4-6 | Anthropic | 88.673 | percent |
| Qwen3.8 Max | #12 | qwen-3-8-max | Alibaba Cloud | 88.2115 | percent |
| GPT-5.4 | #13 | gpt-5-4 | OpenAI | 88.1155 | percent |
| Gemini 3.7 Flash | #14 | gemini-3-7-flash | 87.798 | percent | |
| Muse Spark 1.1 | #15 | muse-spark-1-1 | Meta | 87.73075 | percent |
| Claude Opus 4.7 | #16 | claude-opus-4-7 | Anthropic | 87.19225 | percent |
| Grok 4.5 | #17 | grok-4-5 | xAI | 87.173 | percent |
| DeepSeek V4 Flash 0731 | #18 | deepseek-v4-flash-0731 | DeepSeek | 86.6345 | percent |
| DeepSeek V4 Pro 0813 | #19 | deepseek-v4-pro-0813 | DeepSeek | 85.84125 | percent |
| GPT-5.6 Luna | #20 | gpt-5-6-luna | OpenAI | 85.64425 | percent |
| Gemini 3.6 Flash | #21 | gemini-3-6-flash | 85.149 | percent | |
| Claude Sonnet 4.6 | #22 | claude-sonnet-4-6 | Anthropic | 84.76925 | percent |
| Gemini 3.1 Pro Preview | #23 | gemini-3-1-pro-preview | 84.00475 | percent | |
| Qwen3.7-Max | #24 | qwen-3-7-max | Alibaba Cloud | 83.3365 | percent |
| GPT-5.2 | #25 | gpt-5-2 | OpenAI | 83.2115 | percent |
| Kimi K2.7 Code | #26 | kimi-k2-7-code | Moonshot AI | 82.80775 | percent |
| DeepSeek V4 Pro | #27 | deepseek-v4-pro | DeepSeek | 82.69225 | percent |
| Gemini 3.5 Flash | #28 | gemini-3-5-flash | 82.00475 | percent | |
| GPT-5.4 nano | #29 | gpt-5-4-nano | OpenAI | 81.097 | percent |
| Claude Opus 4.5 | #30 | claude-opus-4-5 | Anthropic | 80.0865 | percent |
| Kimi K2.6 | #31 | kimi-k2-6 | Moonshot AI | 79.375 | percent |
| GLM-5.2 | #32 | glm-5-2 | Z.AI | 78.625 | percent |
| Inkling | #33 | inkling | Thinking Machines Lab | 78.34625 | percent |
| GPT-5.2 Codex | #34 | gpt-5-2-codex | OpenAI | 77.7115 | percent |
| Grok Build 0.1 | #35 | grok-build-0-1 | xAI | 76.37025 | percent |
| Qwen3.6 Plus | #36 | qwen3-6-plus | Alibaba Cloud | 75.827 | percent |
| MiniMax M3 | #37 | minimax-m3 | MiniMax | 74.48075 | percent |
| GPT-5.3-Codex | #38 | gpt-5-3-codex | OpenAI | 72.76 | percent |
| GPT-5.4 mini | #39 | gpt-5-4-mini | OpenAI | 71.32375 | percent |
| Grok 4.3 | #40 | grok-4-3 | xAI | 70.822 | percent |
| DeepSeek V4 Flash | #41 | deepseek-v4-flash | DeepSeek | 70.58175 | percent |
| Qwen3.6 27B | #42 | qwen3-6-27b | Alibaba Cloud | 70.28375 | percent |
| GLM-5.1 | #43 | glm-5-1 | Z.AI | 70.18 | percent |
| Kimi K2.5 | #44 | kimi-k2-5 | Moonshot AI | 69.07 | percent |
| Gemini 3.5 Flash-Lite | #45 | gemini-3-5-flash-lite | 60.1875 | percent |
One best compatible score per canonical product · higher is better
| Rank | Model | Weights | Evidence | Score |
|---|---|---|---|---|
| 01 | GPT-5.6 SolOpenAI | Closed weights | Direct source result | 91.7% |
| 02 | Claude Opus 5Anthropic | Closed weights | Direct source result | 91.2% |
| 03 | Kimi K3Moonshot AI | Open weights | Public reference | 90.7% |
| 04 | GPT-5.6 TerraOpenAI | Closed weights | Direct source result | 90.6% |
| 05 | Grok 4.6xAI | Closed weights | Public reference | 90.5% |
| 06 | Muse Spark 1.2Meta | Open announced | Direct source result | 90.0% |
| 07 | Claude Fable 5Anthropic | Closed weights | Direct source result | 89.7% |
| 08 | GPT-5.5OpenAI | Closed weights | Direct source result | 89.7% |
| 09 | Claude Opus 4.8Anthropic | Closed weights | Direct source result | 89.2% |
| 10 | Claude Sonnet 5Anthropic | Closed weights | Public reference | 88.7% |
Each distinct published measurement is retained. Same-snapshot canonical and configuration projections appear once.
| Published model / configuration | Score | Evidence & protocol | Source & dates |
|---|---|---|---|
gpt-5.6-sol-max (max)GPT-5.6 SolExact identitygpt-5-6-sol-max Canonical product: gpt-5-6-sol | 91.7% | Direct source resultofficial-leaderboardVersion & system2026-06-25 LiveBench owner table — Reasoning partition | LiveBench 2026-06-25 complete owner tableObserved Checked |
claude-opus-5-max-effort (max)Claude Opus 5Exact identityclaude-opus-5-max Canonical product: claude-opus-5 | 91.2% | Direct source resultofficial-leaderboardVersion & system2026-06-25 LiveBench owner table — Reasoning partition | LiveBench 2026-06-25 complete owner tableObserved Checked |
kimi-k3 (reasoning configuration not stated by LiveBench)Kimi K3Exact identitykimi-k3-livebench-2026-06-25-unspecified Canonical product: kimi-k3 | 90.7% | Public referenceofficial-leaderboardVersion & system2026-06-25 LiveBench owner table — Reasoning partition | LiveBench 2026-06-25 complete owner tableObserved Checked |
gpt-5.6-terra-max (max)GPT-5.6 TerraExact identitygpt-5-6-terra-max Canonical product: gpt-5-6-terra | 90.6% | Direct source resultofficial-leaderboardVersion & system2026-06-25 LiveBench owner table — Reasoning partition | LiveBench 2026-06-25 complete owner tableObserved Checked |
grok-4.6 (reasoning configuration not stated by LiveBench)Grok 4.6Exact identitygrok-4-6-livebench-2026-06-25-unspecified Canonical product: grok-4-6 | 90.5% | Public referenceofficial-leaderboardVersion & system2026-06-25 LiveBench owner table — Reasoning partition | LiveBench 2026-06-25 complete owner tableObserved Checked |
muse-spark-1.2-xhigh (xhigh)Muse Spark 1.2Exact identitymuse-spark-1-2-xhigh Canonical product: muse-spark-1-2 | 90.0% | Direct source resultofficial-leaderboardVersion & system2026-06-25 LiveBench owner table — Reasoning partition | LiveBench 2026-06-25 complete owner tableObserved Checked |
gpt-5.5-xhigh (xhigh)GPT-5.5Exact identitygpt-5-5-xhigh Canonical product: gpt-5-5 | 89.7% | Direct source resultofficial-leaderboardVersion & system2026-06-25 LiveBench owner table — Reasoning partition | LiveBench 2026-06-25 complete owner tableObserved Checked |
claude-fable-5-max-effort (max)Claude Fable 5Exact identityclaude-fable-5-max Canonical product: claude-fable-5 | 89.7% | Direct source resultofficial-leaderboardVersion & system2026-06-25 LiveBench owner table — Reasoning partition | LiveBench 2026-06-25 complete owner tableObserved Checked |
claude-opus-4-8-max-effort (max)Claude Opus 4.8Exact identityclaude-opus-4-8-max Canonical product: claude-opus-4-8 | 89.2% | Direct source resultofficial-leaderboardVersion & system2026-06-25 LiveBench owner table — Reasoning partition | LiveBench 2026-06-25 complete owner tableObserved Checked |
claude-sonnet-5-xhigh-effort (xhigh)Claude Sonnet 5Exact identityclaude-sonnet-5-xhigh Canonical product: claude-sonnet-5 | 88.7% | Public referenceofficial-leaderboardVersion & system2026-06-25 LiveBench owner table — Reasoning partition | LiveBench 2026-06-25 complete owner tableObserved Checked |
From result to context
A frequently refreshed benchmark suite designed to limit contamination and cover multiple capabilities.
The summary shows one best compatible source result per canonical product. Versions, effort settings and execution systems remain attached to the underlying records.
Scores from different versions or harnesses may not be interchangeable. The published source rows preserve those distinctions and their original units. A source result is not automatically an input to a current index.
A reviewed track from this benchmark is represented in the following current indices. Each index applies its own protocol and eligibility rules.
Contamination risk: Low. Lifecycle: Rolling.