A task with its own rules.
Measures multi-tool planning and execution across tool-use workloads.
- Organisation
- Version
- May 2026
Agents / Benchmark profile
Measures multi-tool planning and execution across tool-use workloads.
Observed results
20 ranked models · higher is better · labels show rank and score
| Model | Rank | Model ID | Provider | Score | Unit |
|---|---|---|---|---|---|
| Muse Spark 1.1 | #1 | muse-spark-1-1 | Meta | 75.6 | percent |
| Kimi K3 | #2 | kimi-k3 | Moonshot AI | 73.2 | percent |
| GLM-5.3 | #3 | glm-5-3 | Z.AI | 73 | percent |
| Claude Opus 4.8 | #4 | claude-opus-4-8 | Anthropic | 59.9 | percent |
| GPT-5.6 Sol | #5 | gpt-5-6-sol | OpenAI | 58 | percent |
| Gemini 3.5 Flash | #6 | gemini-3-5-flash | 56.5 | percent | |
| GPT-5.5 | #7 | gpt-5-5 | OpenAI | 55.6 | percent |
| GPT-5.4 | #8 | gpt-5-4 | OpenAI | 54.6 | percent |
| GPT-5.6 Luna | #9 | gpt-5-6-luna | OpenAI | 53.4 | percent |
| GPT-5.6 Terra | #10 | gpt-5-6-terra | OpenAI | 53.1 | percent |
| DeepSeek V4 Pro | #11 | deepseek-v4-pro | DeepSeek | 51.8 | percent |
| Kimi K2.6 | #12 | kimi-k2-6 | Moonshot AI | 50 | percent |
| Step 3.7 Flash | #13 | step-3-7-flash | StepFun | 49.5 | percent |
| GLM-5.2 | #14 | glm-5-2 | Z.AI | 48.2 | percent |
| DeepSeek V4 Flash | #15 | deepseek-v4-flash | DeepSeek | 47.8 | percent |
| MiniMax M2.7 | #16 | minimax-m2-7 | MiniMax | 46.3 | percent |
| Claude Opus 4.5 | #17 | claude-opus-4-5 | Anthropic | 43.5 | percent |
| GPT-5.4 mini | #18 | gpt-5-4-mini | OpenAI | 42.9 | percent |
| Qwen3.6 Plus | #19 | qwen3-6-plus | Alibaba Cloud | 39.8 | percent |
| GLM-5 | #20 | glm-5 | Z.AI | 38 | percent |
| Qwen3.5 397B A17B | #21 | qwen3-5-397b | Alibaba Cloud | 36.3 | percent |
| GPT-5.4 nano | #22 | gpt-5-4-nano | OpenAI | 35.5 | percent |
| Kimi K2.5 | #23 | kimi-k2-5 | Moonshot AI | 27.8 | percent |
| Qwen3.6-35B-A3B | #24 | qwen3-6-35b-a3b | Alibaba Cloud | 26.9 | percent |
One best compatible score per canonical product · higher is better
| Rank | Model | Weights | Evidence | Score |
|---|---|---|---|---|
| 01 | Muse Spark 1.1Meta | Closed weights | Direct source result | 75.6% |
| 02 | Kimi K3Moonshot AI | Open weights | Direct source result | 73.2% |
| 03 | GLM-5.3Z.AI | Open announced | Direct source result | 73% |
| 04 | Claude Opus 4.8Anthropic | Closed weights | Source carrier | 59.9% |
| 05 | GPT-5.6 SolOpenAI | Closed weights | Source carrier | 58% |
| 06 | Gemini 3.5 FlashGoogle | Closed weights | Direct source result | 56.5% |
| 07 | GPT-5.5OpenAI | Closed weights | Public reference | 55.6% |
| 08 | GPT-5.4OpenAI | Closed weights | Source carrier | 54.6% |
| 09 | GPT-5.6 LunaOpenAI | Closed weights | Source carrier | 53.4% |
| 10 | GPT-5.6 TerraOpenAI | Closed weights | Source carrier | 53.1% |
Each distinct published measurement is retained. Same-snapshot canonical and configuration projections appear once.
| Published model / configuration | Score | Evidence & protocol | Source & dates |
|---|---|---|---|
Kimi K3 as published by Z.AIKimi K3Exact identitykimi-k3-max Canonical product: kimi-k3 | 76.5% | Relative comparisonprovider-reportedVersion & systemVerified Z.AI GLM-5.3 official comparison table | GLM-5.3: Frontier Coding with Emergent Cyber CapabilitiesObserved Checked |
Claude Opus 4.8 as published by Z.AIClaude Opus 4.8Exact identityclaude-opus-4-8-max Canonical product: claude-opus-4-8 | 76.2% | Relative comparisonprovider-reportedVersion & systemVerified Z.AI GLM-5.3 official comparison table | GLM-5.3: Frontier Coding with Emergent Cyber CapabilitiesObserved Checked |
GPT-5.6 Sol as published by Z.AIGPT-5.6 SolExact identitygpt-5-6-sol-max Canonical product: gpt-5-6-sol | 74.9% | Relative comparisonprovider-reportedVersion & systemVerified Z.AI GLM-5.3 official comparison table | GLM-5.3: Frontier Coding with Emergent Cyber CapabilitiesObserved Checked |
Claude Fable 5 w/ fallback as published by Z.AIClaude Fable 5Exact identityclaude-fable-5-max Canonical product: claude-fable-5 | 74.7% | Relative comparisonprovider-reportedVersion & systemVerified Z.AI GLM-5.3 official comparison table | GLM-5.3: Frontier Coding with Emergent Cyber CapabilitiesObserved Checked |
DeepSeek V4 Pro 0813 as published by Z.AIDeepSeek V4 Pro 0813Exact identitydeepseek-v4-pro-0813-max Canonical product: deepseek-v4-pro-0813 | 74.1% | Relative comparisonprovider-reportedVersion & systemVerified Z.AI GLM-5.3 official comparison table | GLM-5.3: Frontier Coding with Emergent Cyber CapabilitiesObserved Checked |
GLM-5.3 (max)GLM-5.3Exact identityglm-5-3-max Canonical product: glm-5-3 | 73% | Direct source resultprovider-reportedVersion & systemVerified Z.AI GLM-5.3 official comparison table | GLM-5.3: Frontier Coding with Emergent Cyber CapabilitiesObserved Checked |
Qwen3.8 Max as published by Z.AIQwen3.8 MaxExact identityqwen-3-8-max-xhigh Canonical product: qwen-3-8-max | 72.5% | Relative comparisonprovider-reportedVersion & systemVerified Z.AI GLM-5.3 official comparison table | GLM-5.3: Frontier Coding with Emergent Cyber CapabilitiesObserved Checked |
GLM-5.2 as published by Z.AIGLM-5.2Exact identityglm-5-2-max Canonical product: glm-5-2 | 59.9% | Relative comparisonprovider-reportedVersion & systemVerified Z.AI GLM-5.3 official comparison table | GLM-5.3: Frontier Coding with Emergent Cyber CapabilitiesObserved Checked |
Exact BenchLM registry variant Muse Spark 1.1; bulk export does not retain a complete upstream harness configuration.Muse Spark 1.1Exact identitySource label without a registered configuration ID Canonical product: muse-spark-1-1 | 75.6% | Source carriersource-checkedVersion & system2026 Source-native system | BenchLM public datasets — 2026-08-01Observed Checked |
Exact BenchLM registry variant Claude Opus 4.8; bulk export does not retain a complete upstream harness configuration.Claude Opus 4.8Exact identitySource label without a registered configuration ID Canonical product: claude-opus-4-8 | 59.9% | Source carriersource-checkedVersion & system2026 Source-native system | BenchLM public datasets — 2026-08-01Observed Checked |
From result to context
Measures multi-tool planning and execution across tool-use workloads.
The summary shows one best compatible source result per canonical product. Versions, effort settings and execution systems remain attached to the underlying records.
Scores from different versions or harnesses may not be interchangeable. The published source rows preserve those distinctions and their original units. A source result is not automatically an input to a current index.
A reviewed track from this benchmark is represented in the following current indices. Each index applies its own protocol and eligibility rules.
Contamination risk: Unknown. Lifecycle: Active.