A task with its own rules.
The private AutomationBench enterprise-workflow set, kept separate from the public 600-task track and reference-only by default.
- Organisation
- AutomationBench
- Version
- Private set, August 2026
Agents / Benchmark profile
The private AutomationBench enterprise-workflow set, kept separate from the public 600-task track and reference-only by default.
Observed results
4 ranked models · higher is better · labels show rank and score
| Model | Rank | Model ID | Provider | Score | Unit |
|---|---|---|---|---|---|
| Gemini 3.7 Flash | #1 | gemini-3-7-flash | 30.4 | percent | |
| GPT-5.6 Terra | #2 | gpt-5-6-terra | OpenAI | 23.6 | percent |
| Gemini 3.6 Flash | #3 | gemini-3-6-flash | 17 | percent | |
| Claude Sonnet 5 | #4 | claude-sonnet-5 | Anthropic | 10.7 | percent |
One best compatible score per canonical product · higher is better
| Rank | Model | Weights | Evidence | Score |
|---|---|---|---|---|
| 01 | Gemini 3.7 FlashGoogle | Closed weights | Public reference | 30.4% |
| 02 | GPT-5.6 TerraOpenAI | Closed weights | Public reference | 23.6% |
| 03 | Gemini 3.6 FlashGoogle | Closed weights | Public reference | 17% |
| 04 | Claude Sonnet 5Anthropic | Closed weights | Public reference | 10.7% |
Each distinct published measurement is retained. Same-snapshot canonical and configuration projections appear once.
| Published model / configuration | Score | Evidence & protocol | Source & dates |
|---|---|---|---|
Gemini 3.7 FlashGemini 3.7 FlashExact identitygemini-3-7-flash-medium Canonical product: gemini-3-7-flash | 30.4% | Public referenceprovider-reportedVersion & systemPrivate set, August 2026 AutomationBench private set official public leaderboard | Google DeepMind permanent refresh sourceObserved Checked |
GPT-5.6 TerraGPT-5.6 TerraExact identitygpt-5-6-terra-max Canonical product: gpt-5-6-terra | 23.6% | Public referenceprovider-reportedVersion & systemPrivate set, August 2026 AutomationBench private set official public leaderboard | Google DeepMind permanent refresh sourceObserved Checked |
Gemini 3.6 FlashGemini 3.6 FlashExact identitygemini-3-6-flash-high Canonical product: gemini-3-6-flash | 17% | Public referenceprovider-reportedVersion & systemPrivate set, August 2026 AutomationBench private set official public leaderboard | Google DeepMind permanent refresh sourceObserved Checked |
Claude Sonnet 5Claude Sonnet 5Exact identityclaude-sonnet-5-max Canonical product: claude-sonnet-5 | 10.7% | Public referenceprovider-reportedVersion & systemPrivate set, August 2026 AutomationBench private set official public leaderboard | Google DeepMind permanent refresh sourceObserved Checked |
| Model | Result | Execution & attribution | Original source |
|---|---|---|---|
| Qwen3.8 Max (0902)automationBenchPartialScore · exact AA profile b112af07-3bd5-4647-b09f-b23204361bb2 | 56.191percent | AA current published profile; effort label not exportedindependent evaluatorRetained as additional evidence. Existing Core point is unchanged pending an accepted complete own-evidence amendment; incompatible versions remain separate. | Artificial Analysis ↗Publication date not recorded · checked 2026-09-16 |
From result to context
The private AutomationBench enterprise-workflow set, kept separate from the public 600-task track and reference-only by default.
The summary shows one best compatible source result per canonical product. Versions, effort settings and execution systems remain attached to the underlying records.
Scores from different versions or harnesses may not be interchangeable. The published source rows preserve those distinctions and their original units. A source result is not automatically an input to a current index.
A reviewed track from this benchmark is represented in the following current indices. Each index applies its own protocol and eligibility rules.
Contamination risk: Low. Lifecycle: Active.