A task with its own rules.
Tests an agent system's use of tools exposed through Model Context Protocol servers.
- Organisation
- Version
- May 2026
Agents / Benchmark profile
Tests an agent system's use of tools exposed through Model Context Protocol servers.
Observed results
20 ranked models · higher is better · labels show rank and score
| Model | Rank | Model ID | Provider | Score | Unit |
|---|---|---|---|---|---|
| Muse Spark 1.2 | #1 | muse-spark-1-2 | Meta | 90.3 | percent |
| Muse Spark 1.1 | #2 | muse-spark-1-1 | Meta | 88.1 | percent |
| Claude Opus 5 | #3 | claude-opus-5 | Anthropic | 85.8 | percent |
| Kimi K3 | #4 | kimi-k3 | Moonshot AI | 84.2 | percent |
| Gemini 3.5 Flash | #5 | gemini-3-5-flash | 83.6 | percent | |
| Claude Opus 4.8 | #6 | claude-opus-4-8 | Anthropic | 82.2 | percent |
| Inkling-Small | #7 | inkling-small | Thinking Machines Lab | 79.6 | percent |
| Gemini 3.1 Pro Preview | #8 | gemini-3-1-pro-preview | 78.2 | percent | |
| Claude Opus 4.7 | #9 | claude-opus-4-7 | Anthropic | 77.3 | percent |
| GLM-5.2 | #10 | glm-5-2 | Z.AI | 76.8 | percent |
| Qwen3.7-Max | #11 | qwen-3-7-max | Alibaba Cloud | 76.4 | percent |
| Kimi K2.7 Code | #12 | kimi-k2-7-code | Moonshot AI | 76 | percent |
| Muse Glimmer 30B | #13 | muse-glimmer-30b | Meta | 75.5 | percent |
| GPT-5.5 | #14 | gpt-5-5 | OpenAI | 75.3 | percent |
| DeepSeek V4 Pro | #15 | deepseek-v4-pro | DeepSeek | 74.2 | percent |
| MiniMax M3 | #16 | minimax-m3 | MiniMax | 74.2 | percent |
| Inkling | #17 | inkling | Thinking Machines Lab | 74.1 | percent |
| Qwen3.7-Plus | #18 | qwen-3-7-plus | Alibaba Cloud | 73.2 | percent |
| GLM-5.1 | #19 | glm-5-1 | Z.AI | 71.8 | percent |
| GPT-5.4 | #20 | gpt-5-4 | OpenAI | 70.6 | percent |
| DeepSeek V4 Flash | #21 | deepseek-v4-flash | DeepSeek | 69 | percent |
| Qwen3.6-35B-A3B | #22 | qwen3-6-35b-a3b | Alibaba Cloud | 62.8 | percent |
| Solar Open 2 250B | #23 | solar-open2-250b | Upstage | 58.2 | percent |
| GPT-5.4 mini | #24 | gpt-5-4-mini | OpenAI | 57.7 | percent |
| GPT-5.4 nano | #25 | gpt-5-4-nano | OpenAI | 56.1 | percent |
| Kimi K2.6 | #26 | kimi-k2-6 | Moonshot AI | 55.9 | percent |
| Qwen3.6 Plus | #27 | qwen3-6-plus | Alibaba Cloud | 48.2 | percent |
| Qwen3.5 397B A17B | #28 | qwen3-5-397b | Alibaba Cloud | 46.1 | percent |
| Claude Opus 4.5 | #29 | claude-opus-4-5 | Anthropic | 42.3 | percent |
| GLM-5 | #30 | glm-5 | Z.AI | 31.1 | percent |
| Kimi K2.5 | #31 | kimi-k2-5 | Moonshot AI | 29.5 | percent |
One best compatible score per canonical product · higher is better
| Rank | Model | Weights | Evidence | Score |
|---|---|---|---|---|
| 01 | Muse Spark 1.2Meta | Open announced | Direct source result | 90.3% |
| 02 | Muse Spark 1.1Meta | Closed weights | Direct source result | 88.1% |
| 03 | Claude Opus 5Anthropic | Closed weights | Source carrier | 85.8% |
| 04 | Kimi K3Moonshot AI | Open weights | Direct source result | 84.2% |
| 05 | Gemini 3.5 FlashGoogle | Closed weights | Direct source result | 83.6% |
| 06 | Claude Opus 4.8Anthropic | Closed weights | Source carrier | 82.2% |
| 07 | Inkling-SmallThinking Machines Lab | Open weights | Source carrier | 79.6% |
| 08 | Gemini 3.1 Pro PreviewGoogle | Closed weights | Public reference | 78.2% |
| 09 | Claude Opus 4.7Anthropic | Closed weights | Source carrier | 77.3% |
| 10 | GLM-5.2Z.AI | Open weights | Source carrier | 76.8% |
Each distinct published measurement is retained. Same-snapshot canonical and configuration projections appear once.
| Published model / configuration | Score | Evidence & protocol | Source & dates |
|---|---|---|---|
MiMo-V2.5 as published by UpstageMiMo-V2.5Exact identitymimo-v2-5-solar-open2-unspecified Canonical product: mimo-v2-5 | 63.9% | Relative comparisonprovider-reportedVersion & systemstandard Solar Open 2 card English table | Solar Open 2 250B model cardObserved Checked |
DeepSeek V4 Flash max as published by UpstageDeepSeek V4 FlashExact identitydeepseek-v4-flash-max Canonical product: deepseek-v4-flash | 58.2% | Relative comparisonprovider-reportedVersion & systemstandard Solar Open 2 card English table | Solar Open 2 250B model cardObserved Checked |
Solar Open 2 250B (high)Solar Open 2 250BExact identitysolar-open2-250b-high Canonical product: solar-open2-250b | 58.2% | Direct source resultprovider-reportedVersion & systemstandard Solar Open 2 card English table | Solar Open 2 250B model cardObserved Checked |
Solar Open 100B as published by UpstageSolar Open 100B (Reasoning)Exact identitysolar-open-100b-reasoning-high Canonical product: solar-open-100b-reasoning | 34.4% | Relative comparisonprovider-reportedVersion & systemstandard Solar Open 2 card English table | Solar Open 2 250B model cardObserved Checked |
Mistral Medium 3.5 as published by UpstageMistral Medium 3.5Exact identitymistral-medium-3-5-high Canonical product: mistral-medium-3-5 | 30.7% | Relative comparisonprovider-reportedVersion & systemstandard Solar Open 2 card English table | Solar Open 2 250B model cardObserved Checked |
Command A+ as published by UpstageCommand A+Exact identitycommand-a-plus-solar-open2-unspecified Canonical product: command-a-plus | 27.2% | Relative comparisonprovider-reportedVersion & systemstandard Solar Open 2 card English table | Solar Open 2 250B model cardObserved Checked |
Muse Glimmer-30B; high reasoning; temperature=1.0; top_p=0.95; top_k=64Muse Glimmer 30BExact identitymuse-glimmer-30b-high Canonical product: muse-glimmer-30b | 75.5% | Public referenceprovider-reportedVersion & systemPublic / 500 tasks Meta MCP-Atlas evaluation over 500 public tasks, averaged across four runs | Muse Glimmer Evaluation MethodologyObserved Checked |
Muse Spark 1.2 (xhigh)Muse Spark 1.2Exact identitymuse-spark-1-2-xhigh Canonical product: muse-spark-1-2 | 90.3% | Public referenceprovider-reportedVersion & systemMay 2026 Scale AI MCP Atlas comparison as reproduced in Meta's official release chart | Meta Superintelligence Labs permanent refresh sourceObserved Checked |
Muse Spark 1.2 (xhigh) in the Scale AI MCP Atlas harnessMuse Spark 1.2Exact identitymuse-spark-1-2-xhigh Canonical product: muse-spark-1-2 | 90.3% | Public referenceprovider-reportedVersion & systemMay 2026 Scale AI MCP Atlas provider harness | Muse Spark 1.2 Evaluation MethodologyObserved Checked |
Muse Spark 1.1 (reasoning configuration not stated in chart)Muse Spark 1.1Exact identitymuse-spark-1-1-meta-muse-12-release-unspecified Canonical product: muse-spark-1-1 | 88.1% | Relative comparisonprovider-reportedVersion & systemMay 2026 Scale AI MCP Atlas comparison as reproduced in Meta's official release chart | Meta Superintelligence Labs permanent refresh sourceObserved Checked |
From result to context
Tests an agent system's use of tools exposed through Model Context Protocol servers.
The summary shows one best compatible source result per canonical product. Versions, effort settings and execution systems remain attached to the underlying records.
Scores from different versions or harnesses may not be interchangeable. The published source rows preserve those distinctions and their original units. A source result is not automatically an input to a current index.
A reviewed track from this benchmark is represented in the following current indices. Each index applies its own protocol and eligibility rules.
Contamination risk: Unknown. Lifecycle: Active.