A task with its own rules.
A benchmark for agentic professional workflows with verifiable success criteria, reporting pass rates and partial scores for model plus agent-harness rows.
- Organisation
- UC Berkeley RDI
- Version
- 2026
Knowledge / Benchmark profile
A benchmark for agentic professional workflows with verifiable success criteria, reporting pass rates and partial scores for model plus agent-harness rows.
Observed results
16 ranked models · higher is better · labels show rank and score
| Model | Rank | Model ID | Provider | Score | Unit |
|---|---|---|---|---|---|
| Qwen3.8-Flash-Next | #1 | qwen-3-8-flash-next | Alibaba Cloud | 51.2 | percent |
| Qwen3.8-27B | #2 | qwen-3-8-27b | Alibaba Cloud | 42.9 | percent |
| Claude Sonnet 5 | #3 | claude-sonnet-5 | Anthropic | 33.3 | percent |
| GLM-5.3 | #4 | glm-5-3 | Z.AI | 28.5 | percent |
| GPT-5.6 Terra | #5 | gpt-5-6-terra | OpenAI | 28 | percent |
| Kimi K3 | #6 | kimi-k3 | Moonshot AI | 27.6 | percent |
| DeepSeek V4 Flash Vision Exp | #7 | deepseek-v4-flash-vision-exp | DeepSeek | 27.3 | percent |
| Gemini 3.7 Flash | #8 | gemini-3-7-flash | 26.3 | percent | |
| GLM-5.3-Flash | #9 | glm-5-3-flash | Z.AI | 26.3 | percent |
| Claude Opus 4.8 | #10 | claude-opus-4-8 | Anthropic | 25.7 | percent |
| DeepSeek V4 Pro 0813 | #11 | deepseek-v4-pro-0813 | DeepSeek | 25.7 | percent |
| DeepSeek V4 Flash | #12 | deepseek-v4-flash | DeepSeek | 25.2 | percent |
| DeepSeek V4 Flash 0731 | #13 | deepseek-v4-flash-0731 | DeepSeek | 25.2 | percent |
| Gemini 3.6 Flash | #14 | gemini-3-6-flash | 24.2 | percent | |
| GLM-5.2 | #15 | glm-5-2 | Z.AI | 23.8 | percent |
| DeepSeek V4 Pro | #16 | deepseek-v4-pro | DeepSeek | 16.5 | percent |
One best compatible score per canonical product · higher is better
| Rank | Model | Weights | Evidence | Score |
|---|---|---|---|---|
| 01 | Qwen3.8-Flash-NextAlibaba Cloud | Open weights | Public reference | 51.2% |
| 02 | Qwen3.8-27BAlibaba Cloud | Open weights | Public reference | 42.9% |
| 03 | Claude Sonnet 5Anthropic | Closed weights | Public reference | 33.3% |
| 04 | GLM-5.3Z.AI | Open announced | Public reference | 28.5% |
| 05 | GPT-5.6 TerraOpenAI | Closed weights | Public reference | 28% |
| 06 | Kimi K3Moonshot AI | Open weights | Public reference | 27.6% |
| 07 | DeepSeek V4 Flash Vision ExpDeepSeek | Not documented | Public reference | 27.3% |
| 08 | Gemini 3.7 FlashGoogle | Closed weights | Public reference | 26.3% |
| 09 | GLM-5.3-FlashZ.AI | Open weights | Public reference | 26.3% |
| 10 | Claude Opus 4.8Anthropic | Closed weights | Public reference | 25.7% |
Each distinct published measurement is retained. Same-snapshot canonical and configuration projections appear once.
| Published model / configuration | Score | Evidence & protocol | Source & dates |
|---|---|---|---|
Qwen3.8-Flash-Next (xhigh default thinking configuration)Qwen3.8-Flash-NextExact identityqwen-3-8-flash-next-xhigh Canonical product: qwen-3-8-flash-next | 51.2% | Public referencesource-checkedVersion & systemScore system:qwen3-8-flash-next:ale-score | Qwen3.8-Flash-Next current official launch pageObserved Checked |
Qwen3.8-27B (xhigh default thinking configuration) as published in Qwen's Qwen3.8-Flash-Next comparison tableQwen3.8-27BExact identityqwen-3-8-27b-xhigh Canonical product: qwen-3-8-27b | 42.9% | Relative comparisonprovider-reportedVersion & systemScore system:qwen3-8-flash-next-comparison:cell:language:ale-score:1 | Qwen3.8-Flash-Next launch and official provider evaluationsObserved Checked |
Qwen3.7-Plus (provider-published configuration) as published in Qwen's Qwen3.8-Flash-Next comparison tableQwen3.7-PlusExact identityqwen-3-7-plus-unspecified Canonical product: qwen-3-7-plus | 33.6% | Relative comparisonprovider-reportedVersion & systemScore system:qwen3-8-flash-next-comparison:cell:language:ale-score:2 | Qwen3.8-Flash-Next launch and official provider evaluationsObserved Checked |
GLM-5.3-Flash (max effort, documented default)GLM-5.3-FlashExact identityglm-5-3-flash-max Canonical product: glm-5-3-flash | 26.3% | Public referencesource-checkedVersion & system2026 Claude Code | GLM-5.3-Flash individual evaluationsObserved Checked |
DeepSeek-V4-Flash-0731 (provider-published configuration) as published in Qwen's Qwen3.8-Flash-Next comparison tableDeepSeek V4 Flash 0731Exact identitydeepseek-v4-flash-0731-deepseek-0813-release-unspecified Canonical product: deepseek-v4-flash-0731 | 25.2% | Source carrierprovider-reportedVersion & systemPass@1 system:qwen3-8-flash-next-comparison:cell:language:ale-pass1:3 | Qwen3.8-Flash-Next launch and official provider evaluationsObserved Checked |
Qwen3.8-Flash-Next (xhigh default thinking configuration)Qwen3.8-Flash-NextExact identityqwen-3-8-flash-next-xhigh Canonical product: qwen-3-8-flash-next | 24.3% | Public referencesource-checkedVersion & systemPass@1 system:qwen3-8-flash-next:ale-pass1 | Qwen3.8-Flash-Next current official launch pageObserved Checked |
Qwen3.8-27B (xhigh default thinking configuration) as published in Qwen's Qwen3.8-Flash-Next comparison tableQwen3.8-27BExact identityqwen-3-8-27b-xhigh Canonical product: qwen-3-8-27b | 20.4% | Relative comparisonprovider-reportedVersion & systemPass@1 system:qwen3-8-flash-next-comparison:cell:language:ale-pass1:1 | Qwen3.8-Flash-Next launch and official provider evaluationsObserved Checked |
Qwen3.7-Plus (provider-published configuration) as published in Qwen's Qwen3.8-Flash-Next comparison tableQwen3.7-PlusExact identityqwen-3-7-plus-unspecified Canonical product: qwen-3-7-plus | 13.2% | Relative comparisonprovider-reportedVersion & systemPass@1 system:qwen3-8-flash-next-comparison:cell:language:ale-pass1:2 | Qwen3.8-Flash-Next launch and official provider evaluationsObserved Checked |
Provider chart does not specify an exact evaluation systemDeepSeek V4 Flash Vision ExpExact identitySource label without a registered configuration ID Canonical product: deepseek-v4-flash-vision-exp | 27.3% | Public referenceprovider-reportedVersion & system2026 Source-native system | DeepSeek-V4-Flash-Vision-Exp release and provider evaluationObserved Checked |
Qwen3.8-27B (xhigh)Qwen3.8-27BExact identityqwen-3-8-27b-xhigh Canonical product: qwen-3-8-27b | 42.9% | Public referenceprovider-reportedVersion & systemScore Qwen3.8-27B official text table | Qwen3.8-27B official model cardObserved Checked |
From result to context
A benchmark for agentic professional workflows with verifiable success criteria, reporting pass rates and partial scores for model plus agent-harness rows.
The summary shows one best compatible source result per canonical product. Versions, effort settings and execution systems remain attached to the underlying records.
Scores from different versions or harnesses may not be interchangeable. The published source rows preserve those distinctions and their original units. A source result is not automatically an input to a current index.
This catalogue entry preserves available source evidence. Current indices admit only their specifically reviewed tracks and configurations.
Contamination risk: Unknown. Lifecycle: Archived.