A task with its own rules.
The source-native pairwise Elo track for Artificial Analysis' GDPval-AA v2 evaluation, kept separate from the normalized percentage track to prevent unit mixing.
- Organisation
- Artificial Analysis / OpenAI
- Version
- v2
Agents / Benchmark profile
The source-native pairwise Elo track for Artificial Analysis' GDPval-AA v2 evaluation, kept separate from the normalized percentage track to prevent unit mixing.
Observed results
12 ranked models · higher is better · labels show rank and score
| Model | Rank | Model ID | Provider | Score | Unit |
|---|---|---|---|---|---|
| GLM-5.3 | #1 | glm-5-3 | Z.AI | 1769 | Elo |
| Grok 4.6 | #2 | grok-4-6 | xAI | 1753 | Elo |
| Claude Fable 5 | #3 | claude-fable-5 | Anthropic | 1741 | Elo |
| GPT-5.6 Sol | #4 | gpt-5-6-sol | OpenAI | 1728 | Elo |
| Muse Spark 1.2 | #5 | muse-spark-1-2 | Meta | 1631 | Elo |
| Claude Sonnet 5 | #6 | claude-sonnet-5 | Anthropic | 1598 | Elo |
| GPT-5.6 Terra | #7 | gpt-5-6-terra | OpenAI | 1578 | Elo |
| Grok 4.5 | #8 | grok-4-5 | xAI | 1526 | Elo |
| Gemini 3.7 Flash | #9 | gemini-3-7-flash | 1525 | Elo | |
| Gemini 3.6 Flash | #10 | gemini-3-6-flash | 1422 | Elo | |
| Solar Open 2 250B | #11 | solar-open2-250b | Upstage | 1128 | Elo |
| NVIDIA Nemotron 3.5 Lightning 30B-A3B | #12 | nemotron-3-5-lightning-30b-a3b | NVIDIA | 832 | Elo |
One best compatible score per canonical product · higher is better
| Rank | Model | Weights | Evidence | Score |
|---|---|---|---|---|
| 01 | GLM-5.3Z.AI | Open announced | Public reference | 1769 Elo |
| 02 | Grok 4.6xAI | Closed weights | Public reference | 1753 Elo |
| 03 | Claude Fable 5Anthropic | Closed weights | Public reference | 1741 Elo |
| 04 | GPT-5.6 SolOpenAI | Closed weights | Public reference | 1728 Elo |
| 05 | Muse Spark 1.2Meta | Open announced | Public reference | 1631 Elo |
| 06 | Claude Sonnet 5Anthropic | Closed weights | Public reference | 1598 Elo |
| 07 | GPT-5.6 TerraOpenAI | Closed weights | Public reference | 1578 Elo |
| 08 | Grok 4.5xAI | Closed weights | Public reference | 1526 Elo |
| 09 | Gemini 3.7 FlashGoogle | Closed weights | Public reference | 1525 Elo |
| 10 | Gemini 3.6 FlashGoogle | Closed weights | Public reference | 1422 Elo |
Each distinct published measurement is retained. Same-snapshot canonical and configuration projections appear once.
| Published model / configuration | Score | Evidence & protocol | Source & dates |
|---|---|---|---|
DeepSeek V4 Flash max as published by UpstageDeepSeek V4 FlashExact identitydeepseek-v4-flash-max Canonical product: deepseek-v4-flash | 1187 Elo | Relative comparisonprovider-reportedVersion & systemv2 Solar Open 2 card English table | Solar Open 2 250B model cardObserved Checked |
MiMo-V2.5 as published by UpstageMiMo-V2.5Exact identitymimo-v2-5-solar-open2-unspecified Canonical product: mimo-v2-5 | 1145 Elo | Relative comparisonprovider-reportedVersion & systemv2 Solar Open 2 card English table | Solar Open 2 250B model cardObserved Checked |
Solar Open 2 250B (high)Solar Open 2 250BExact identitysolar-open2-250b-high Canonical product: solar-open2-250b | 1128 Elo | Public referenceprovider-reportedVersion & systemv2 Solar Open 2 card English table | Solar Open 2 250B model cardObserved Checked |
Mistral Medium 3.5 as published by UpstageMistral Medium 3.5Exact identitymistral-medium-3-5-high Canonical product: mistral-medium-3-5 | 929 Elo | Relative comparisonprovider-reportedVersion & systemv2 Solar Open 2 card English table | Solar Open 2 250B model cardObserved Checked |
Command A+ as published by UpstageCommand A+Exact identitycommand-a-plus-solar-open2-unspecified Canonical product: command-a-plus | 712 Elo | Relative comparisonprovider-reportedVersion & systemv2 Solar Open 2 card English table | Solar Open 2 250B model cardObserved Checked |
GLM-5.3 (max)GLM-5.3Exact identityglm-5-3-max Canonical product: glm-5-3 | 1769 Elo | Public referenceprovider-reportedVersion & systemv2 Z.AI GLM-5.3 official comparison table | GLM-5.3: Frontier Coding with Emergent Cyber CapabilitiesObserved Checked |
Claude Fable 5 w/ fallback as published by Z.AIClaude Fable 5Exact identityclaude-fable-5-max Canonical product: claude-fable-5 | 1743 Elo | Relative comparisonprovider-reportedVersion & systemv2 Z.AI GLM-5.3 official comparison table | GLM-5.3: Frontier Coding with Emergent Cyber CapabilitiesObserved Checked |
Qwen3.8 Max as published by Z.AIQwen3.8 MaxExact identityqwen-3-8-max-xhigh Canonical product: qwen-3-8-max | 1739 Elo | Relative comparisonprovider-reportedVersion & systemv2 Z.AI GLM-5.3 official comparison table | GLM-5.3: Frontier Coding with Emergent Cyber CapabilitiesObserved Checked |
GPT-5.6 Sol as published by Z.AIGPT-5.6 SolExact identitygpt-5-6-sol-max Canonical product: gpt-5-6-sol | 1730 Elo | Relative comparisonprovider-reportedVersion & systemv2 Z.AI GLM-5.3 official comparison table | GLM-5.3: Frontier Coding with Emergent Cyber CapabilitiesObserved Checked |
Kimi K3 as published by Z.AIKimi K3Exact identitykimi-k3-max Canonical product: kimi-k3 | 1682 Elo | Relative comparisonprovider-reportedVersion & systemv2 Z.AI GLM-5.3 official comparison table | GLM-5.3: Frontier Coding with Emergent Cyber CapabilitiesObserved Checked |
| Model | Result | Execution & attribution | Original source |
|---|---|---|---|
| Qwen3.8 Max (0902)gdpval · exact AA profile b112af07-3bd5-4647-b09f-b23204361bb2 | 1,688.96native score | AA current published profile; effort label not exportedindependent evaluatorRetained as additional evidence. Existing Core point is unchanged pending an accepted complete own-evidence amendment; incompatible versions remain separate. | Artificial Analysis ↗Publication date not recorded · checked 2026-09-16 |
From result to context
The source-native pairwise Elo track for Artificial Analysis' GDPval-AA v2 evaluation, kept separate from the normalized percentage track to prevent unit mixing.
The summary shows one best compatible source result per canonical product. Versions, effort settings and execution systems remain attached to the underlying records.
Scores from different versions or harnesses may not be interchangeable. The published source rows preserve those distinctions and their original units. A source result is not automatically an input to a current index.
A reviewed track from this benchmark is represented in the following current indices. Each index applies its own protocol and eligibility rules.
Contamination risk: Low. Lifecycle: Active.