Registry pulse
Evidence at a glance.
Families with scores
72/84
Sourced records
1987
Weighted families
49
Capability areas
8
Visual summaries are derived from the same repository records shown in each benchmark table.
Evaluation directory
Open any family for a clean per-benchmark leaderboard — top models, full score table, about, FAQ and compare links. 84 families in the registry.
84 of 84 benchmark families · 1987 sourced result records · 49 weighted families
Registry pulse
Families with scores
72/84
Sourced records
1987
Weighted families
49
Capability areas
8
Visual summaries are derived from the same repository records shown in each benchmark table.
Category coverage
Capability category
Artificial Analysis · Current / rolling
Independently evaluated automation and agentic workflow benchmark from Artificial Analysis.
#1 · Moonshot AI
Kimi K3
52.7%
#2 · xAI
Grok 4.5
51.4%
#3 · OpenAI
GPT-5.6 Sol
51.2%
The top three are 1: Kimi K3, 52.7%; 2: Grok 4.5, 51.4%; 3: GPT-5.6 Sol, 51.2%. Open the benchmark for interactive charts and an accessible table.
Artificial Analysis · 2026
Artificial Analysis private agentic knowledge-work evaluation with Elo over rubric pass rate, analytical quality, and presentation quality on business deliverables.
#1 · Meta
Muse Spark 1.1
86.3 elo-proxy
#2 · Moonshot AI
Kimi K3
7.6 elo-proxy
The top three are 1: Muse Spark 1.1, 86.3 elo-proxy; 2: Kimi K3, 7.6 elo-proxy. Open the benchmark for interactive charts and an accessible table.
THUDM · Current / rolling
A multi-environment benchmark for evaluating language models as interactive agents.
Artificial Analysis / Mercor · Current / rolling
Long-horizon professional multi-application agent benchmark implemented independently by Artificial Analysis.
#1 · Google
Gemini 3.5 Flash
47.1%
#2 · OpenAI
GPT-5.5
37.7%
#3 · Moonshot AI
Kimi K3
37.6%
The top three are 1: Gemini 3.5 Flash, 47.1%; 2: GPT-5.5, 37.7%; 3: Kimi K3, 37.6%. Open the benchmark for interactive charts and an accessible table.
BenchLM · bench-align-v5.1
Category prior from the BenchLM overall leaderboard agentic score used to estimate missing agentic coverage.
#1 · Anthropic
Claude Mythos 5
77.1%
#2 · Anthropic
Claude Fable 5
76.8%
#3 · OpenAI
GPT-5.6 Sol
75.5%
The top three are 1: Claude Mythos 5, 77.1%; 2: Claude Fable 5, 76.8%; 3: GPT-5.6 Sol, 75.5%. Open the benchmark for interactive charts and an accessible table.
Berkeley Gorilla · V3
An evaluation suite for function selection, argument construction and tool-use behaviour.
#1 · Alibaba Cloud
Qwen3.7-Max
75%
#2 · Alibaba Cloud
Qwen3.7-Plus
72.9%
The top three are 1: Qwen3.7-Max, 75%; 2: Qwen3.7-Plus, 72.9%. Open the benchmark for interactive charts and an accessible table.
OpenAI · Current / rolling
A benchmark for browsing agents that must locate difficult-to-find information on the web.
#1 · OpenAI
GPT-5.6 Sol
92.2%
#2 · Moonshot AI
Kimi K3
91.2%
#3 · OpenAI
GPT-5.5 Pro
90.1%
The top three are 1: GPT-5.6 Sol, 92.2%; 2: Kimi K3, 91.2%; 3: GPT-5.5 Pro, 90.1%. Open the benchmark for interactive charts and an accessible table.
Stanford / Cybench authors · Current / rolling
Professional-level Capture-the-Flag cybersecurity agent benchmark measuring unguided end-to-end task solve rates.
#1 · Anthropic
Claude Mythos 5
100%
#2 · Anthropic
Claude Opus 4.7
96%
#3 · Anthropic
Claude Opus 4.6
93%
The top three are 1: Claude Mythos 5, 100%; 2: Claude Opus 4.7, 96%; 3: Claude Opus 4.6, 93%. Open the benchmark for interactive charts and an accessible table.
Artificial Analysis / ServiceNow · Current / rolling
Stateful multi-step enterprise workflow agent benchmark graded on final database state across business domains.
#1 · Anthropic
Claude Fable 5
51.1%
#2 · Google
Gemini 3.5 Flash
50.1%
#3 · OpenAI
GPT-5.5
46.6%
The top three are 1: Claude Fable 5, 51.1%; 2: Gemini 3.5 Flash, 50.1%; 3: GPT-5.5, 46.6%. Open the benchmark for interactive charts and an accessible table.
ExploitBench authors · 2
A cybersecurity-agent benchmark measuring progress from reaching vulnerable code through exploit primitives and code execution.
#1 · OpenAI
GPT-5.5
34%
#2 · Anthropic
Claude Opus 4.7
27%
The top three are 1: GPT-5.5, 34%; 2: Claude Opus 4.7, 27%. Open the benchmark for interactive charts and an accessible table.
ExploitGym authors · 2026-05
A realistic cybersecurity benchmark for evaluating whether agents can turn known software vulnerabilities into concrete attacks.
#1 · OpenAI
GPT-5.6 Sol
33.7%
#2 · OpenAI
GPT-5.6 Terra
23.2%
#3 · Anthropic
Claude Mythos 5
17.5%
The top three are 1: GPT-5.6 Sol, 33.7%; 2: GPT-5.6 Terra, 23.2%; 3: Claude Mythos 5, 17.5%. Open the benchmark for interactive charts and an accessible table.
GAIA authors · Current / rolling
A benchmark for general assistants that must reason, browse, use tools and combine modalities.
#1 · Anthropic
Claude Mythos 5
52.3%
#2 · Anthropic
Claude Fable 5
52.3%
#3 · OpenAI
GPT-5.4
48.2%
The top three are 1: Claude Mythos 5, 52.3%; 2: Claude Fable 5, 52.3%; 3: GPT-5.4, 48.2%. Open the benchmark for interactive charts and an accessible table.
Artificial Analysis / OpenAI · v2
Artificial Analysis agentic evaluation of OpenAI GDPval real-world economic work tasks across occupations, scored with pairwise Elo then normalized.
#1 · Anthropic
Claude Fable 5
63.0%
#2 · OpenAI
GPT-5.6 Sol
62.4%
#3 · Moonshot AI
Kimi K3
59.2%
The top three are 1: Claude Fable 5, 63.0%; 2: GPT-5.6 Sol, 62.4%; 3: Kimi K3, 59.2%. Open the benchmark for interactive charts and an accessible table.
Artificial Analysis / Harvey · Current / rolling
Legal agent benchmark on private Harvey legal work tasks graded criterion-by-criterion in an agent sandbox.
#1 · Moonshot AI
Kimi K3
26.7%
#2 · Anthropic
Claude Fable 5
14.2%
#3 · xAI
Grok 4.5
13.3%
The top three are 1: Kimi K3, 26.7%; 2: Claude Fable 5, 14.2%; 3: Grok 4.5, 13.3%. Open the benchmark for interactive charts and an accessible table.
Artificial Analysis / IBM · Current / rolling
Artificial Analysis implementation of IBM ITBench for Kubernetes incident root-cause analysis from offline snapshots.
#1 · OpenAI
GPT-5.6 Sol
56.2%
#2 · OpenAI
GPT-5.6 Terra
51.0%
#3 · Anthropic
Claude Opus 4.7
46.7%
The top three are 1: GPT-5.6 Sol, 56.2%; 2: GPT-5.6 Terra, 51.0%; 3: Claude Opus 4.7, 46.7%. Open the benchmark for interactive charts and an accessible table.
Long-Horizon-Terminal-Bench authors · 2026-07
A 46-task terminal benchmark for long-running agent work with dense subtask grading across nine practical domains.
#1 · OpenAI
GPT-5.6 Sol
65.9%
#2 · Anthropic
Claude Fable 5
62.9%
#3 · OpenAI
GPT-5.5
60.6%
The top three are 1: GPT-5.6 Sol, 65.9%; 2: Claude Fable 5, 62.9%; 3: GPT-5.5, 60.6%. Open the benchmark for interactive charts and an accessible table.
Google · May 2026
Tests an agent system's use of tools exposed through Model Context Protocol servers.
#1 · Meta
Muse Spark 1.1
88.1%
#2 · Moonshot AI
Kimi K3
84.2%
#3 · Google
Gemini 3.5 Flash
83.6%
The top three are 1: Muse Spark 1.1, 88.1%; 2: Kimi K3, 84.2%; 3: Gemini 3.5 Flash, 83.6%. Open the benchmark for interactive charts and an accessible table.
OfficeQA authors · Current / rolling
Document and spreadsheet workplace QA benchmark measuring office-style multi-file reasoning and calculation.
Single qualifying result
Kimi K3
Moonshot AI
OSWorld authors · Current / rolling
A benchmark for multimodal agents performing tasks in real computer operating systems.
#1 · Anthropic
Claude Opus 4.5
66.3%
#2 · OpenAI
GPT-5.6 Sol
62.6%
#3 · OpenAI
GPT-5.6 Terra
50.2%
The top three are 1: Claude Opus 4.5, 66.3%; 2: GPT-5.6 Sol, 62.6%; 3: GPT-5.6 Terra, 50.2%. Open the benchmark for interactive charts and an accessible table.
OSWorld · Verified
Evaluates computer-use agents in verified desktop interaction tasks.
#1 · Anthropic
Claude Fable 5
85%
#2 · Anthropic
Claude Mythos 5
85%
#3 · Anthropic
Claude Opus 4.8
83.4%
The top three are 1: Claude Fable 5, 85%; 2: Claude Mythos 5, 85%; 3: Claude Opus 4.8, 83.4%. Open the benchmark for interactive charts and an accessible table.
Terminal-Bench team · 2.1
A benchmark for agents performing practical tasks in reproducible terminal environments.
#1 · OpenAI
GPT-5.6 Sol
91.9%
#2 · Moonshot AI
Kimi K3
88.3%
#3 · OpenAI
GPT-5.6 Terra
88.0%
The top three are 1: GPT-5.6 Sol, 91.9%; 2: Kimi K3, 88.3%; 3: GPT-5.6 Terra, 88.0%. Open the benchmark for interactive charts and an accessible table.
Laude Institute / Artificial Analysis · Current / rolling
Harder agentic terminal evaluation suite spanning software engineering, systems administration, and data processing tasks.
#1 · Moonshot AI
Kimi K3
85%
#2 · OpenAI
GPT-5.6 Sol
65.9%
#3 · Anthropic
Claude Fable 5
62.9%
The top three are 1: Kimi K3, 85%; 2: GPT-5.6 Sol, 65.9%; 3: Claude Fable 5, 62.9%. Open the benchmark for interactive charts and an accessible table.
Google · May 2026
Measures multi-tool planning and execution across tool-use workloads.
#1 · Meta
Muse Spark 1.1
75.6%
#2 · Moonshot AI
Kimi K3
73.2%
#3 · Anthropic
Claude Opus 4.8
59.9%
The top three are 1: Muse Spark 1.1, 75.6%; 2: Kimi K3, 73.2%; 3: Claude Opus 4.8, 59.9%. Open the benchmark for interactive charts and an accessible table.
WebArena authors · Current / rolling
A benchmark of autonomous agents completing realistic tasks on self-hosted web applications.
Single qualifying result
Muse Spark 1.1
Meta
WebVoyager authors · Current / rolling
End-to-end web browser agent benchmark evaluating multi-step navigation and task completion on live sites.
Sierra Research · Current / rolling
A benchmark for tool-using language agents interacting with simulated user and business environments.
#1 · Z.AI
GLM-5.2
99.1%
#2 · Anthropic
Claude Fable 5
98.5%
#3 · Z.AI
GLM-5V-Turbo
98.5%
The top three are 1: GLM-5.2, 99.1%; 2: Claude Fable 5, 98.5%; 3: GLM-5V-Turbo, 98.5%. Open the benchmark for interactive charts and an accessible table.
Sierra / Artificial Analysis · 2
Dual-control conversational agent benchmark for telecom support where agent and user must coordinate tool actions.
#1 · Z.AI
GLM-5.2
99.1%
#2 · Anthropic
Claude Fable 5
98.5%
#3 · Z.AI
GLM-5-Turbo
98.5%
The top three are 1: GLM-5.2, 99.1%; 2: Claude Fable 5, 98.5%; 3: GLM-5-Turbo, 98.5%. Open the benchmark for interactive charts and an accessible table.
Sierra / Artificial Analysis · 3
Fintech customer-support agent benchmark requiring knowledge-base navigation and multi-step tool use on banking workflows.
#1 · Moonshot AI
Kimi K3
33.4%
#2 · OpenAI
GPT-5.6 Sol
33.0%
#3 · xAI
Grok 4.5
32.6%
The top three are 1: Kimi K3, 33.4%; 2: GPT-5.6 Sol, 33.0%; 3: Grok 4.5, 32.6%. Open the benchmark for interactive charts and an accessible table.
Capability category
BenchLM · bench-align-v5.1
Category prior from the BenchLM overall leaderboard coding score used to estimate missing coding coverage.
#1 · Anthropic
Claude Mythos 5
82.0%
#2 · Anthropic
Claude Fable 5
81.7%
#3 · OpenAI
GPT-5.6 Sol
74.1%
The top three are 1: Claude Mythos 5, 82.0%; 2: Claude Fable 5, 81.7%; 3: GPT-5.6 Sol, 74.1%. Open the benchmark for interactive charts and an accessible table.
BigCodeBench authors · Current / rolling
A benchmark for practical code generation involving diverse libraries and complex instructions.
CRUXEval authors · Current / rolling
A code-reasoning benchmark based on predicting program inputs and outputs.
Cursor · 3.2
Cursor's first-party benchmark for ambiguous multi-file coding-agent tasks from real Cursor sessions (v3.2 public snapshot).
#1 · Anthropic
Claude Fable 5
70.5%
#2 · OpenAI
GPT-5.6 Sol
67.2%
#3 · xAI
Grok 4.5
66.7%
The top three are 1: Claude Fable 5, 70.5%; 2: GPT-5.6 Sol, 67.2%; 3: Grok 4.5, 66.7%. Open the benchmark for interactive charts and an accessible table.
DataCurve · v1.1
Contamination-resistant software engineering repair tasks across real repositories under a shared agent harness.
Single qualifying result
Kimi K3
Moonshot AI
FrontierSWE · 2026-07
Extremely difficult implementation, performance, and research software tasks scored with pairwise dominance.
Single qualifying result
Kimi K3
Moonshot AI
Moonshot AI · 2.0
Moonshot internal end-to-end coding tasks used in the Kimi K3 launch table; retained reference-only.
Single qualifying result
Kimi K3
Moonshot AI
LiveCodeBench team · Current / rolling
A continuously updated benchmark for code generation, execution, self-repair and related tasks.
#1 · Google
Gemini 3 Pro
91.7%
#2 · Alibaba Cloud
Qwen3.7-Max
91.6%
#3 · Alibaba Cloud
Qwen3.7-Plus
89.6%
The top three are 1: Gemini 3 Pro, 91.7%; 2: Qwen3.7-Max, 91.6%; 3: Qwen3.7-Plus, 89.6%. Open the benchmark for interactive charts and an accessible table.
Multi-SWE-Bench · Current / rolling
Multi-language software engineering repository evaluation extending SWE-bench style tasks.
Single qualifying result
MiniMax M2.7
MiniMax
PostTrainBench · public
Autonomous post-training of small language models under fixed compute and time budgets.
Single qualifying result
Kimi K3
Moonshot AI
ProgramBench · public
Rebuild command-line programs from binary behaviour and documentation without source code.
Single qualifying result
Kimi K3
Moonshot AI
SciCode authors · Current / rolling
A benchmark for generating scientific-computing code from expert-authored specifications.
#1 · Anthropic
Claude Fable 5
60.2%
#2 · Google
Gemini 3.1 Pro Preview
58.9%
#3 · Moonshot AI
Kimi K3
58.7%
The top three are 1: Claude Fable 5, 60.2%; 2: Gemini 3.1 Pro Preview, 58.9%; 3: Kimi K3, 58.7%. Open the benchmark for interactive charts and an accessible table.
Abundant AI · v1.1
Multi-hour whole-project software engineering tasks with hidden checks and exploit scanning.
Single qualifying result
Kimi K3
Moonshot AI
Scale AI · v1
Public single-attempt software engineering tasks drawn from repositories outside the original SWE-bench set.
#1 · Anthropic
Claude Mythos 5
80.3%
#2 · Anthropic
Claude Fable 5
80%
#3 · Anthropic
Claude Opus 4.8
69.2%
The top three are 1: Claude Mythos 5, 80.3%; 2: Claude Fable 5, 80%; 3: Claude Opus 4.8, 69.2%. Open the benchmark for interactive charts and an accessible table.
SWE-bench authors · Verified
A human-validated subset of real GitHub software-engineering issue resolution tasks.
#1 · Anthropic
Claude Mythos 5
95.5%
#2 · Anthropic
Claude Fable 5
95%
#3 · Anthropic
Claude Opus 4.8
88.6%
The top three are 1: Claude Mythos 5, 95.5%; 2: Claude Fable 5, 95%; 3: Claude Opus 4.8, 88.6%. Open the benchmark for interactive charts and an accessible table.
SWE-Rebench · Current / rolling
Software engineering repository tasks designed as a refreshed SWE-bench style evaluation.
#1 · Anthropic
Claude Opus 4.6
65.3%
#2 · Z.AI
GLM-5
62.8%
#3 · Z.AI
GLM-5.1
62.7%
The top three are 1: Claude Opus 4.6, 65.3%; 2: GLM-5, 62.8%; 3: GLM-5.1, 62.7%. Open the benchmark for interactive charts and an accessible table.
Capability category
LMSYS Org · Current / rolling
An automated pairwise evaluation set derived from challenging user prompts.
Single qualifying result
Mistral Large 3
Mistral AI
IFBench · Current / rolling
Instruction-following evaluation measuring constraint and format reliability.
#1 · MiniMax
MiniMax M3
82.9%
#2 · NVIDIA
Nemotron 3 Ultra
81.4%
#3 · xAI
Grok 4.3
81.3%
The top three are 1: MiniMax M3, 82.9%; 2: Nemotron 3 Ultra, 81.4%; 3: Grok 4.3, 81.3%. Open the benchmark for interactive charts and an accessible table.
Google Research · Current / rolling
An evaluation of whether language models follow verifiable natural-language instructions.
#1 · Alibaba Cloud
Qwen3.7-Plus
94.6%
#2 · Alibaba Cloud
Qwen3.7-Max
94.3%
#3 · Alibaba Cloud
Qwen3.6 Plus
94.3%
The top three are 1: Qwen3.7-Plus, 94.6%; 2: Qwen3.7-Max, 94.3%; 3: Qwen3.6 Plus, 94.3%. Open the benchmark for interactive charts and an accessible table.
WildBench authors · Current / rolling
A benchmark based on challenging real-world user conversations and automated evaluation.
Capability category
THUDM · 2
A long-context reasoning benchmark built around demanding real-world tasks.
#1 · Alibaba Cloud
Qwen3.7-Plus
91.7%
#2 · Alibaba Cloud
Qwen3.7-Max
90.4%
#3 · Google
Gemini 3.5 Flash
77.3%
The top three are 1: Qwen3.7-Plus, 91.7%; 2: Qwen3.7-Max, 90.4%; 3: Gemini 3.5 Flash, 77.3%. Open the benchmark for interactive charts and an accessible table.
Capability category
MAA / public contest evals · Current / rolling
American Invitational Mathematics Examination 2025 contest problems used as a frontier math evaluation.
#1 · OpenAI
GPT-5.2
99%
#2 · OpenAI
GPT-5 Codex
98.7%
#3 · Xiaomi
MiMo-V2-Flash
96.3%
The top three are 1: GPT-5.2, 99%; 2: GPT-5 Codex, 98.7%; 3: MiMo-V2-Flash, 96.3%. Open the benchmark for interactive charts and an accessible table.
MAA / public contest evals · Current / rolling
American Invitational Mathematics Examination 2026 contest problems used as a frontier math evaluation.
#1 · Z.AI
GLM-5.2
99.2%
#2 · Z.AI
GLM-5
95.8%
#3 · Moonshot AI
Kimi K2.5
95.8%
The top three are 1: GLM-5.2, 99.2%; 2: GLM-5, 95.8%; 3: Kimi K2.5, 95.8%. Open the benchmark for interactive charts and an accessible table.
BenchLM · bench-align-v5.1
Category prior from the BenchLM math score used to estimate missing mathematics coverage.
#1 · OpenAI
GPT-5.6 Sol
95.3%
#2 · OpenAI
GPT-5.6 Terra
95.3%
#3 · OpenAI
GPT-5.6 Luna
95.3%
The top three are 1: GPT-5.6 Sol, 95.3%; 2: GPT-5.6 Terra, 95.3%; 3: GPT-5.6 Luna, 95.3%. Open the benchmark for interactive charts and an accessible table.
Epoch AI · Current / rolling
A benchmark of difficult, expert-authored mathematics problems designed for frontier systems.
#1 · OpenAI
GPT-5.6 Sol
89%
#2 · OpenAI
GPT-5.6 Terra
84.9%
#3 · OpenAI
GPT-5.6 Luna
78.6%
The top three are 1: GPT-5.6 Sol, 89%; 2: GPT-5.6 Terra, 84.9%; 3: GPT-5.6 Luna, 78.6%. Open the benchmark for interactive charts and an accessible table.
HMMT · Current / rolling
Harvard-MIT Mathematics Tournament February 2026 problems used as a contest math evaluation.
#1 · Alibaba Cloud
Qwen3.7-Max
97.1%
#2 · Alibaba Cloud
Qwen3.7-Plus
92.9%
#3 · Z.AI
GLM-5.2
92.5%
The top three are 1: Qwen3.7-Max, 97.1%; 2: Qwen3.7-Plus, 92.9%; 3: GLM-5.2, 92.5%. Open the benchmark for interactive charts and an accessible table.
OpenAI / Hendrycks · Current / rolling
A 500-problem competition-math subset from the MATH dataset spanning algebra, geometry, number theory, and related domains.
#1 · OpenAI
GPT-5
99.4%
#2 · OpenAI
o3
99.2%
#3 · xAI
Grok 3 mini
99.2%
The top three are 1: GPT-5, 99.4%; 2: o3, 99.2%; 3: Grok 3 mini, 99.2%. Open the benchmark for interactive charts and an accessible table.
MAA / public contest evals · Current / rolling
United States of America Mathematical Olympiad 2026 problems used as an expert math evaluation.
#1 · Anthropic
Claude Mythos 5
99.8%
#2 · Anthropic
Claude Opus 4.8
96.7%
#3 · MiniMax
MiniMax M3
85.7%
The top three are 1: Claude Mythos 5, 99.8%; 2: Claude Opus 4.8, 96.7%; 3: MiniMax M3, 85.7%. Open the benchmark for interactive charts and an accessible table.
Capability category
BenchLM · bench-align-v5.1
Category prior from the BenchLM multimodal score used to estimate missing multimodal coverage.
#1 · Google
Gemini 3 Pro Deep Think
95%
#2 · xAI
Grok 4.1
93.1%
#3 · Anthropic
Claude Mythos 5
92.8%
The top three are 1: Gemini 3 Pro Deep Think, 95%; 2: Grok 4.1, 93.1%; 3: Claude Mythos 5, 92.8%. Open the benchmark for interactive charts and an accessible table.
Google · v2
Measures agentic spatial reasoning over blueprint-style visual information.
#1 · Anthropic
Claude Fable 5
38.6%
#2 · OpenAI
GPT-5.5
36.2%
#3 · Google
Gemini 3.5 Flash
33.6%
The top three are 1: Claude Fable 5, 38.6%; 2: GPT-5.5, 36.2%; 3: Gemini 3.5 Flash, 33.6%. Open the benchmark for interactive charts and an accessible table.
ChartQA authors · Current / rolling
A visual question-answering benchmark focused on charts and data visualisations.
CharXiv · May 2026
Measures reasoning over charts in scientific papers.
#1 · Anthropic
Claude Mythos 5
93.5%
#2 · Moonshot AI
Kimi K3
91.3%
#3 · Anthropic
Claude Opus 4.8
89.9%
The top three are 1: Claude Mythos 5, 93.5%; 2: Kimi K3, 91.3%; 3: Claude Opus 4.8, 89.9%. Open the benchmark for interactive charts and an accessible table.
DocVQA organisers · Current / rolling
A benchmark family for visual question answering over document images.
MathVista authors · Current / rolling
A benchmark for mathematical reasoning over diagrams, charts and other visual inputs.
MMMU authors · Current / rolling
A multi-discipline multimodal understanding and reasoning benchmark for expert-level tasks.
Single qualifying result
Qwen3.6 Plus
Alibaba Cloud
MMMU-Pro authors · Current / rolling
A more robust extension of MMMU intended to reduce shortcut-based answering.
#1 · Google
Gemini 3.1 Pro Preview
83.9%
#2 · Google
Gemini 3.5 Flash
83.9%
#3 · OpenAI
GPT-5.6 Sol
83.4%
The top three are 1: Gemini 3.1 Pro Preview, 83.9%; 2: Gemini 3.5 Flash, 83.9%; 3: GPT-5.6 Sol, 83.4%. Open the benchmark for interactive charts and an accessible table.
OCRBench authors · 2
Multimodal OCR and document understanding benchmark covering text recognition in complex real-world images.
Capability category
ARC Prize Foundation · 2
A second-generation abstract reasoning benchmark based on novel grid transformation tasks.
#1 · OpenAI
GPT-5.5
85%
#2 · Google
Gemini 3.1 Pro Preview
77.1%
#3 · Google
Gemini 3.5 Flash
72.1%
The top three are 1: GPT-5.5, 85%; 2: Gemini 3.1 Pro Preview, 77.1%; 3: Gemini 3.5 Flash, 72.1%. Open the benchmark for interactive charts and an accessible table.
ARC Prize Foundation · 3
An interactive reasoning benchmark testing exploration, world modelling, goal discovery, planning and adaptive execution in novel environments.
#1 · OpenAI
GPT-5.6 Sol
7.8%
#2 · OpenAI
GPT-5.6 Terra
0.8%
#3 · OpenAI
GPT-5.6 Luna
0.2%
The top three are 1: GPT-5.6 Sol, 7.8%; 2: GPT-5.6 Terra, 0.8%; 3: GPT-5.6 Luna, 0.2%. Open the benchmark for interactive charts and an accessible table.
BenchLM · bench-align-v5.1
Category prior from the BenchLM overall leaderboard reasoning score used to estimate missing reasoning coverage.
#1 · xAI
Grok 4.1
90.8%
#2 · Google
Gemini 3 Pro Deep Think
88.3%
#3 · Anthropic
Claude Opus 4.6
87.8%
The top three are 1: Grok 4.1, 90.8%; 2: Gemini 3 Pro Deep Think, 88.3%; 3: Claude Opus 4.6, 87.8%. Open the benchmark for interactive charts and an accessible table.
CritPt authors · Current / rolling
Research-level physics reasoning benchmark with composite challenges designed to stress frontier scientific reasoning.
#1 · OpenAI
GPT-5.6 Sol
32.3%
#2 · OpenAI
GPT-5.5 Pro
30.6%
#3 · OpenAI
GPT-5.6 Terra
30%
The top three are 1: GPT-5.6 Sol, 32.3%; 2: GPT-5.5 Pro, 30.6%; 3: GPT-5.6 Terra, 30%. Open the benchmark for interactive charts and an accessible table.
GPQA authors · Current / rolling
The Diamond subset of a graduate-level, expert-written question-answering benchmark.
#1 · OpenAI
GPT-5.6 Sol
94.6%
#2 · Google
Gemini 3.1 Pro Preview
94.1%
#3 · Anthropic
Claude Mythos 5
94.1%
The top three are 1: GPT-5.6 Sol, 94.6%; 2: Gemini 3.1 Pro Preview, 94.1%; 3: Claude Mythos 5, 94.1%. Open the benchmark for interactive charts and an accessible table.
Center for AI Safety and Scale AI · Current / rolling
A broad expert-level academic benchmark covering difficult questions across many domains.
#1 · Anthropic
Claude Mythos 5
64.5%
#2 · Meta
Muse Spark 1.1
62.1%
#3 · Anthropic
Claude Opus 4.8
57.9%
The top three are 1: Claude Mythos 5, 64.5%; 2: Muse Spark 1.1, 62.1%; 3: Claude Opus 4.8, 57.9%. Open the benchmark for interactive charts and an accessible table.
LiveBench team · Current / rolling
A frequently refreshed benchmark suite designed to limit contamination and cover multiple capabilities.
#1 · OpenAI
GPT-5.5
80.7%
#2 · OpenAI
GPT-5.4
80.3%
#3 · Google
Gemini 3.1 Pro Preview
79.9%
The top three are 1: GPT-5.5, 80.7%; 2: GPT-5.4, 80.3%; 3: Gemini 3.1 Pro Preview, 79.9%. Open the benchmark for interactive charts and an accessible table.
MuSR authors · Current / rolling
Multi-step soft reasoning benchmark requiring long narrative understanding and structured logical deduction.
SuperGPQA authors · Current / rolling
Graduate-level Google-proof Q&A expansion beyond GPQA Diamond for harder knowledge-reasoning items.
#1 · Anthropic
Claude Opus 4.6
95%
#2 · Anthropic
Claude Sonnet 4.6
95%
#3 · Alibaba Cloud
Qwen3.7-Max
73.6%
The top three are 1: Claude Opus 4.6, 95%; 2: Claude Sonnet 4.6, 95%; 3: Qwen3.7-Max, 73.6%. Open the benchmark for interactive charts and an accessible table.
Capability category
Artificial Analysis · Current / rolling
Artificial Analysis long-context reasoning benchmark measuring extraction and synthesis over documents from roughly 10k to 100k tokens.
#1 · OpenAI
GPT-5.2 Codex
75.7%
#2 · OpenAI
GPT-5
75.6%
#3 · OpenAI
GPT-5.1
75%
The top three are 1: GPT-5.2 Codex, 75.7%; 2: GPT-5, 75.6%; 3: GPT-5.1, 75%. Open the benchmark for interactive charts and an accessible table.
Artificial Analysis · Current / rolling
Artificial Analysis knowledge and hallucination benchmark measuring factual recall across economically relevant domains.
#1 · Meta
Muse Spark 1.1
40.6%
#2 · Anthropic
Claude Fable 5
40.1 index
#3 · Google
Gemini 3.1 Pro Preview
32.9 index
The top three are 1: Muse Spark 1.1, 40.6%; 2: Claude Fable 5, 40.1 index; 3: Gemini 3.1 Pro Preview, 32.9 index. Open the benchmark for interactive charts and an accessible table.
BenchLM · bench-align-v5.1
Category prior from the BenchLM knowledge score used to estimate missing research coverage.
#1 · OpenAI
GPT-5.4
96.8%
#2 · Anthropic
Claude Opus 4.6
91.5%
#3 · xAI
Grok 4.1
90.1%
The top three are 1: GPT-5.4, 96.8%; 2: Claude Opus 4.6, 91.5%; 3: Grok 4.1, 90.1%. Open the benchmark for interactive charts and an accessible table.
Google · v2
Measures multi-step financial research and evidence synthesis by an agent system.
#1 · Google
Gemini 3.5 Flash
57.9%
#2 · Meta
Muse Spark 1.1
57.2%
#3 · OpenAI
GPT-5.5
51.8%
The top three are 1: Gemini 3.5 Flash, 57.9%; 2: Muse Spark 1.1, 57.2%; 3: GPT-5.5, 51.8%. Open the benchmark for interactive charts and an accessible table.
Cohere Labs · Current / rolling
Lightweight multilingual MMLU-style knowledge benchmark spanning diverse languages and cultural contexts.
OpenAI · Current / rolling
OpenAI HealthBench hard subset evaluating model helpfulness and safety on challenging healthcare conversations.
TIGER-Lab · Current / rolling
A more challenging multiple-choice knowledge and reasoning benchmark derived from MMLU.
#1 · Google
Gemini 3 Pro
89.8%
#2 · Alibaba Cloud
Qwen3.7-Max
89.6%
#3 · Anthropic
Claude Opus 4.5
89.5%
The top three are 1: Gemini 3 Pro, 89.8%; 2: Qwen3.7-Max, 89.6%; 3: Claude Opus 4.5, 89.5%. Open the benchmark for interactive charts and an accessible table.
OpenAI · 2
Multi-needle long-context retrieval and reasoning benchmark stressing multi-document evidence gathering.
Single qualifying result
Muse Spark 1.1
Meta
OpenAI · Current / rolling
An evaluation of agents attempting to replicate machine-learning research from published papers.
#1 · Moonshot AI
Kimi K2.5
63.5%
#2 · MiniMax
MiniMax M3
52.6%
The top three are 1: Kimi K2.5, 63.5%; 2: MiniMax M3, 52.6%. Open the benchmark for interactive charts and an accessible table.
OpenAI · Current / rolling
A short-answer factuality benchmark designed to measure correctness on fact-seeking questions.
#1 · DeepSeek
DeepSeek V4 Pro
45%
#2 · DeepSeek
DeepSeek V4 Flash
23.1%
The top three are 1: DeepSeek V4 Pro, 45%; 2: DeepSeek V4 Flash, 23.1%. Open the benchmark for interactive charts and an accessible table.