A task with its own rules.
An independently evaluated enterprise-operations benchmark from Artificial Analysis.
- Organisation
- Artificial Analysis
- Version
- 2026
Agents / Benchmark profile
An independently evaluated enterprise-operations benchmark from Artificial Analysis.
Observed results
19 ranked models · higher is better · labels show rank and score
| Model | Rank | Model ID | Provider | Score | Unit |
|---|---|---|---|---|---|
| Claude Fable 5 | #1 | claude-fable-5 | Anthropic | 51.1 | percent |
| Gemini 3.5 Flash | #2 | gemini-3-5-flash | 50.1 | percent | |
| Muse Spark 1.1 | #3 | muse-spark-1-1 | Meta | 47.2 | percent |
| GPT-5.5 | #4 | gpt-5-5 | OpenAI | 46.6 | percent |
| Kimi K3 | #5 | kimi-k3 | Moonshot AI | 45.3 | percent |
| Qwen3.7-Max | #6 | qwen-3-7-max | Alibaba Cloud | 45 | percent |
| Claude Sonnet 5 | #7 | claude-sonnet-5 | Anthropic | 44.7 | percent |
| Claude Opus 4.8 | #8 | claude-opus-4-8 | Anthropic | 44 | percent |
| GPT-5.6 Sol | #9 | gpt-5-6-sol | OpenAI | 42.9 | percent |
| GLM-5.2 | #10 | glm-5-2 | Z.AI | 42.7 | percent |
| Grok 4.5 | #11 | grok-4-5 | xAI | 40.8 | percent |
| DeepSeek V4 Pro | #12 | deepseek-v4-pro | DeepSeek | 40.4 | percent |
| DeepSeek V4 Flash | #13 | deepseek-v4-flash | DeepSeek | 39.6 | percent |
| Inkling | #14 | inkling | Thinking Machines Lab | 38.1 | percent |
| Mistral Medium 3.5 128B | #15 | mistral-medium-3-5-128b | Mistral AI | 33.7 | percent |
| MiniMax M3 | #16 | minimax-m3 | MiniMax | 32.1 | percent |
| Nemotron 3 Ultra | #17 | nemotron-3-ultra | NVIDIA | 28.9 | percent |
| Gemma 4 31B | #18 | gemma-4-31b | 28.3 | percent | |
| GPT-OSS 120B | #19 | gpt-oss-120b | OpenAI | 25.5 | percent |
One best compatible score per canonical product · higher is better
| Rank | Model | Weights | Evidence | Score |
|---|---|---|---|---|
| 01 | Claude Fable 5Anthropic | Closed weights | Source carrier | 51.1% |
| 02 | Gemini 3.5 FlashGoogle | Closed weights | Source carrier | 50.1% |
| 03 | Muse Spark 1.1Meta | Closed weights | Source carrier | 47.2% |
| 04 | GPT-5.5OpenAI | Closed weights | Source carrier | 46.6% |
| 05 | Kimi K3Moonshot AI | Open weights | Source carrier | 45.3% |
| 06 | Qwen3.7-MaxAlibaba Cloud | Closed weights | Source carrier | 45% |
| 07 | Claude Sonnet 5Anthropic | Closed weights | Source carrier | 44.7% |
| 08 | Claude Opus 4.8Anthropic | Closed weights | Source carrier | 44% |
| 09 | GPT-5.6 SolOpenAI | Closed weights | Source carrier | 42.9% |
| 10 | GLM-5.2Z.AI | Open weights | Source carrier | 42.7% |
Each distinct published measurement is retained. Same-snapshot canonical and configuration projections appear once.
| Published model / configuration | Score | Evidence & protocol | Source & dates |
|---|---|---|---|
Exact BenchLM registry variant Claude Fable 5; bulk export does not retain a complete upstream harness configuration.Claude Fable 5Exact identitySource label without a registered configuration ID Canonical product: claude-fable-5 | 51.1% | Source carriersource-checkedVersion & system2026 Source-native system | BenchLM public datasets — 2026-08-01Observed Checked |
Exact BenchLM registry variant Gemini 3.5 Flash; bulk export does not retain a complete upstream harness configuration.Gemini 3.5 FlashExact identitySource label without a registered configuration ID Canonical product: gemini-3-5-flash | 50.1% | Source carriersource-checkedVersion & system2026 Source-native system | BenchLM public datasets — 2026-08-01Observed Checked |
Exact BenchLM registry variant Muse Spark 1.1; bulk export does not retain a complete upstream harness configuration.Muse Spark 1.1Exact identitySource label without a registered configuration ID Canonical product: muse-spark-1-1 | 47.2% | Source carriersource-checkedVersion & system2026 Source-native system | BenchLM public datasets — 2026-08-01Observed Checked |
Exact BenchLM registry variant GPT-5.5; bulk export does not retain a complete upstream harness configuration.GPT-5.5Exact identitySource label without a registered configuration ID Canonical product: gpt-5-5 | 46.6% | Source carriersource-checkedVersion & system2026 Source-native system | BenchLM public datasets — 2026-08-01Observed Checked |
Exact BenchLM registry variant Kimi K3; bulk export does not retain a complete upstream harness configuration.Kimi K3Exact identitySource label without a registered configuration ID Canonical product: kimi-k3 | 45.3% | Source carriersource-checkedVersion & system2026 Source-native system | BenchLM public datasets — 2026-08-01Observed Checked |
Exact BenchLM registry variant Qwen3.7 Max; bulk export does not retain a complete upstream harness configuration.Qwen3.7-MaxExact identitySource label without a registered configuration ID Canonical product: qwen-3-7-max | 45% | Source carriersource-checkedVersion & system2026 Source-native system | BenchLM public datasets — 2026-08-01Observed Checked |
Exact BenchLM registry variant Claude Sonnet 5; bulk export does not retain a complete upstream harness configuration.Claude Sonnet 5Exact identitySource label without a registered configuration ID Canonical product: claude-sonnet-5 | 44.7% | Source carriersource-checkedVersion & system2026 Source-native system | BenchLM public datasets — 2026-08-01Observed Checked |
Exact BenchLM registry variant Claude Opus 4.8; bulk export does not retain a complete upstream harness configuration.Claude Opus 4.8Exact identitySource label without a registered configuration ID Canonical product: claude-opus-4-8 | 44% | Source carriersource-checkedVersion & system2026 Source-native system | BenchLM public datasets — 2026-08-01Observed Checked |
Exact BenchLM registry variant GPT-5.6 Sol; bulk export does not retain a complete upstream harness configuration.GPT-5.6 SolExact identitySource label without a registered configuration ID Canonical product: gpt-5-6-sol | 42.9% | Source carriersource-checkedVersion & system2026 Source-native system | BenchLM public datasets — 2026-08-01Observed Checked |
Exact BenchLM registry variant GLM-5.2; bulk export does not retain a complete upstream harness configuration.GLM-5.2Exact identitySource label without a registered configuration ID Canonical product: glm-5-2 | 42.7% | Source carriersource-checkedVersion & system2026 Source-native system | BenchLM public datasets — 2026-08-01Observed Checked |
From result to context
An independently evaluated enterprise-operations benchmark from Artificial Analysis.
The summary shows one best compatible source result per canonical product. Versions, effort settings and execution systems remain attached to the underlying records.
Scores from different versions or harnesses may not be interchangeable. The published source rows preserve those distinctions and their original units. A source result is not automatically an input to a current index.
This catalogue entry preserves available source evidence. Current indices admit only their specifically reviewed tracks and configurations.
Contamination risk: Unknown. Lifecycle: Active.