A task with its own rules.
Long-horizon professional multi-application agent benchmark implemented independently by Artificial Analysis.
- Organisation
- Artificial Analysis / Mercor
- Version
- Current / rolling
Agents / Benchmark profile
Long-horizon professional multi-application agent benchmark implemented independently by Artificial Analysis.
Observed results
20 ranked models · higher is better · labels show rank and score
| Model | Rank | Model ID | Provider | Score | Unit |
|---|---|---|---|---|---|
| Claude Fable 5 | #1 | claude-fable-5 | Anthropic | 59.2 | percent |
| Grok 4.6 | #2 | grok-4-6 | xAI | 57.5 | percent |
| GPT-5.6 Sol | #3 | gpt-5-6-sol | OpenAI | 56.7 | percent |
| Gemini 3.5 Flash | #4 | gemini-3-5-flash | 47.1 | percent | |
| Grok 4.5 | #5 | grok-4-5 | xAI | 47.1 | percent |
| Kimi K3 | #6 | kimi-k3 | Moonshot AI | 41.3 | percent |
| GPT-5.6 Terra | #7 | gpt-5-6-terra | OpenAI | 38.9 | percent |
| GPT-5.5 | #8 | gpt-5-5 | OpenAI | 37.7 | percent |
| GPT-5.6 Luna | #9 | gpt-5-6-luna | OpenAI | 35.8 | percent |
| GLM-5.2 | #10 | glm-5-2 | Z.AI | 33.7021 | percent |
| GPT-5.4 | #11 | gpt-5-4 | OpenAI | 33.3 | percent |
| Claude Opus 4.6 | #12 | claude-opus-4-6 | Anthropic | 33.0383 | percent |
| Gemini 3.1 Pro Preview | #13 | gemini-3-1-pro-preview | 32.0059 | percent | |
| Kimi K2.6 | #14 | kimi-k2-6 | Moonshot AI | 28.5 | percent |
| GPT-5.4 mini | #15 | gpt-5-4-mini | OpenAI | 28.2 | percent |
| Claude Sonnet 4.6 | #16 | claude-sonnet-4-6 | Anthropic | 28.0236 | percent |
| Gemini 3 Flash | #17 | gemini-3-flash | 27.7286 | percent | |
| Hy3 | #18 | hy3 | Tencent | 25.6 | percent |
| GPT-5.4 nano | #19 | gpt-5-4-nano | OpenAI | 24.9263 | percent |
| DeepSeek V4 Pro | #20 | deepseek-v4-pro | DeepSeek | 24.3 | percent |
| Qwen3.7-Plus | #21 | qwen-3-7-plus | Alibaba Cloud | 22.4189 | percent |
| Grok 4.3 | #22 | grok-4-3 | xAI | 17.0354 | percent |
| Solar Open 2 250B | #23 | solar-open2-250b | Upstage | 16.6 | percent |
| Qwen3.5 397B A17B | #24 | qwen3-5-397b | Alibaba Cloud | 15.3392 | percent |
| Step 3.7 Flash | #25 | step-3-7-flash | StepFun | 14.8 | percent |
| DeepSeek V3.2 | #26 | deepseek-v3-2 | DeepSeek | 14.528 | percent |
| GLM-5 | #27 | glm-5 | Z.AI | 14.5 | percent |
| Gemini 3.1 Flash-Lite | #28 | gemini-3-1-flash-lite | 12.2 | percent | |
| Kimi K2.5 | #29 | kimi-k2-5 | Moonshot AI | 11.5044 | percent |
| MiniMax M2.7 | #30 | minimax-m2-7 | MiniMax | 10.6195 | percent |
| GPT-OSS 120B | #31 | gpt-oss-120b | OpenAI | 3.1 | percent |
| MiMo-V2.5-Pro | #32 | mimo-v2-5-pro | Xiaomi | 2.4336 | percent |
| Nemotron 3 Super | #33 | nemotron-3-super | NVIDIA | 1.8437 | percent |
| GPT-OSS 20B | #34 | gpt-oss-20b | OpenAI | 0.7 | percent |
One best compatible score per canonical product · higher is better
| Rank | Model | Weights | Evidence | Score |
|---|---|---|---|---|
| 01 | Claude Fable 5Anthropic | Closed weights | Relative comparison | 59.2% |
| 02 | Grok 4.6xAI | Closed weights | Direct source result | 57.5% |
| 03 | GPT-5.6 SolOpenAI | Closed weights | Relative comparison | 56.7% |
| 04 | Gemini 3.5 FlashGoogle | Closed weights | Source carrier | 47.1% |
| 05 | Grok 4.5xAI | Closed weights | Relative comparison | 47.1% |
| 06 | Kimi K3Moonshot AI | Open weights | Source carrier | 41.3% |
| 07 | GPT-5.6 TerraOpenAI | Closed weights | Source carrier | 38.9% |
| 08 | GPT-5.5OpenAI | Closed weights | Source carrier | 37.7% |
| 09 | GPT-5.6 LunaOpenAI | Closed weights | Source carrier | 35.8% |
| 10 | GLM-5.2Z.AI | Open weights | Direct source result | 33.7% |
Each distinct published measurement is retained. Same-snapshot canonical and configuration projections appear once.
| Published model / configuration | Score | Evidence & protocol | Source & dates |
|---|---|---|---|
Solar Open 2 250B (high)Solar Open 2 250BExact identitysolar-open2-250b-high Canonical product: solar-open2-250b | 16.6% | Direct source resultprovider-reportedVersion & systemstandard Solar Open 2 card English table | Solar Open 2 250B model cardObserved Checked |
MiMo-V2.5 as published by UpstageMiMo-V2.5Exact identitymimo-v2-5-solar-open2-unspecified Canonical product: mimo-v2-5 | 13.4% | Relative comparisonprovider-reportedVersion & systemstandard Solar Open 2 card English table | Solar Open 2 250B model cardObserved Checked |
DeepSeek V4 Flash max as published by UpstageDeepSeek V4 FlashExact identitydeepseek-v4-flash-max Canonical product: deepseek-v4-flash | 13.2% | Relative comparisonprovider-reportedVersion & systemstandard Solar Open 2 card English table | Solar Open 2 250B model cardObserved Checked |
Mistral Medium 3.5 as published by UpstageMistral Medium 3.5Exact identitymistral-medium-3-5-high Canonical product: mistral-medium-3-5 | 6.1% | Relative comparisonprovider-reportedVersion & systemstandard Solar Open 2 card English table | Solar Open 2 250B model cardObserved Checked |
Solar Open 100B as published by UpstageSolar Open 100B (Reasoning)Exact identitysolar-open-100b-reasoning-high Canonical product: solar-open-100b-reasoning | 2.4% | Relative comparisonprovider-reportedVersion & systemstandard Solar Open 2 card English table | Solar Open 2 250B model cardObserved Checked |
Command A+ as published by UpstageCommand A+Exact identitycommand-a-plus-solar-open2-unspecified Canonical product: command-a-plus | 1.6% | Relative comparisonprovider-reportedVersion & systemstandard Solar Open 2 card English table | Solar Open 2 250B model cardObserved Checked |
Fable 5 MaxClaude Fable 5Exact identityclaude-fable-5-max Canonical product: claude-fable-5 | 59.2% | Relative comparisonprovider-reportedVersion & systemCurrent xAI Grok 4.6 official release comparison table | Grok 4.6Observed Checked |
Grok 4.6 HighGrok 4.6Exact identitygrok-4-6-high Canonical product: grok-4-6 | 57.5% | Direct source resultprovider-reportedVersion & systemCurrent xAI Grok 4.6 official release comparison table | Grok 4.6Observed Checked |
GPT-5.6 Sol MaxGPT-5.6 SolExact identitygpt-5-6-sol-max Canonical product: gpt-5-6-sol | 56.7% | Relative comparisonprovider-reportedVersion & systemCurrent xAI Grok 4.6 official release comparison table | Grok 4.6Observed Checked |
Grok 4.5 HighGrok 4.5Exact identitygrok-4-5-aa-2-high Canonical product: grok-4-5 | 47.1% | Relative comparisonprovider-reportedVersion & systemCurrent xAI Grok 4.6 official release comparison table | Grok 4.6Observed Checked |
From result to context
Long-horizon professional multi-application agent benchmark implemented independently by Artificial Analysis.
The summary shows one best compatible source result per canonical product. Versions, effort settings and execution systems remain attached to the underlying records.
Scores from different versions or harnesses may not be interchangeable. The published source rows preserve those distinctions and their original units. A source result is not automatically an input to a current index.
This catalogue entry preserves available source evidence. Current indices admit only their specifically reviewed tracks and configurations.
Contamination risk: Unknown. Lifecycle: Active.