A task with its own rules.
Cybersecurity agent benchmark focused on finding and validating software vulnerabilities under constrained tool environments.
- Organisation
- UC Berkeley SunBlaze
- Version
- Current / rolling
Agents / Benchmark profile
Cybersecurity agent benchmark focused on finding and validating software vulnerabilities under constrained tool environments.
Observed results
20 ranked models · higher is better · labels show rank and score
| Model | Rank | Model ID | Provider | Score | Unit |
|---|---|---|---|---|---|
| Fugu Cyber | #1 | sakana-fugu-cyber | Sakana AI | 86.9 | percent |
| GLM-5.3 | #2 | glm-5-3 | Z.AI | 84.5 | percent |
| GPT-5.6 Sol | #3 | gpt-5-6-sol | OpenAI | 84.5 | percent |
| DeepSeek V4 Pro 0813 | #4 | deepseek-v4-pro-0813 | DeepSeek | 83.3 | percent |
| Gemini 3.5 Flash Cyber | #5 | gemini-3-5-flash-cyber | 83.2 | percent | |
| Claude Fable 5 | #6 | claude-fable-5 | Anthropic | 83.1 | percent |
| GPT-5.5 | #7 | gpt-5-5 | OpenAI | 81.8 | percent |
| GPT-5.6 Terra | #8 | gpt-5-6-terra | OpenAI | 81.8 | percent |
| Kimi K3 | #9 | kimi-k3 | Moonshot AI | 80 | percent |
| GPT-5.4 | #10 | gpt-5-4 | OpenAI | 79 | percent |
| Claude Opus 4.8 | #11 | claude-opus-4-8 | Anthropic | 78.3 | percent |
| GPT-5.6 Luna | #12 | gpt-5-6-luna | OpenAI | 77.9 | percent |
| DeepSeek V4 Flash | #13 | deepseek-v4-flash | DeepSeek | 76.7 | percent |
| DeepSeek V4 Flash 0731 | #14 | deepseek-v4-flash-0731 | DeepSeek | 76.7 | percent |
| DeepSeek V4 Flash Vision Exp | #15 | deepseek-v4-flash-vision-exp | DeepSeek | 75.3 | percent |
| Claude Opus 4.7 | #16 | claude-opus-4-7 | Anthropic | 73.1 | percent |
| GLM-5.1 | #17 | glm-5-1 | Z.AI | 68.7 | percent |
| Claude Opus 4.6 | #18 | claude-opus-4-6 | Anthropic | 66.6 | percent |
| Claude Sonnet 4.6 | #19 | claude-sonnet-4-6 | Anthropic | 65.2 | percent |
| Muse Spark 1.1 | #20 | muse-spark-1-1 | Meta | 59 | percent |
| DeepSeek V4 Pro | #21 | deepseek-v4-pro | DeepSeek | 52.7 | percent |
| Claude Opus 4.5 | #22 | claude-opus-4-5 | Anthropic | 50.6 | percent |
| Muse Spark | #23 | muse-spark | Meta | 43.5 | percent |
| GLM-5 | #24 | glm-5 | Z.AI | 43.2 | percent |
One best compatible score per canonical product · higher is better
| Rank | Model | Weights | Evidence | Score |
|---|---|---|---|---|
| 01 | Fugu CyberSakana AI | Closed weights | Source carrier | 86.9% |
| 02 | GLM-5.3Z.AI | Open announced | Public reference | 84.5% |
| 03 | GPT-5.6 SolOpenAI | Closed weights | Source carrier | 84.5% |
| 04 | DeepSeek V4 Pro 0813DeepSeek | Open weights | Public reference | 83.3% |
| 05 | Gemini 3.5 Flash CyberGoogle | Closed weights | Source carrier | 83.2% |
| 06 | Claude Fable 5Anthropic | Closed weights | Public reference | 83.1% |
| 07 | GPT-5.5OpenAI | Closed weights | Source carrier | 81.8% |
| 08 | GPT-5.6 TerraOpenAI | Closed weights | Source carrier | 81.8% |
| 09 | Kimi K3Moonshot AI | Open weights | Public reference | 80% |
| 10 | GPT-5.4OpenAI | Closed weights | Source carrier | 79% |
Each distinct published measurement is retained. Same-snapshot canonical and configuration projections appear once.
| Published model / configuration | Score | Evidence & protocol | Source & dates |
|---|---|---|---|
DeepSeek Harness Minimal Mode; max effort; top_p=0.95; temperature=1.0DeepSeek V4 Flash Vision ExpExact identitydeepseek-v4-flash-vision-exp-max-harness Canonical product: deepseek-v4-flash-vision-exp | 75.3% | Public referenceprovider-reportedVersion & systemstandard DeepSeek Harness Minimal Mode | DeepSeek-V4-Flash-Vision-Exp release and provider evaluationObserved Checked |
GLM-5.3 (max)GLM-5.3Exact identityglm-5-3-max Canonical product: glm-5-3 | 84.5% | Public referenceprovider-reportedVersion & systemrolling Z.AI GLM-5.3 official comparison table | GLM-5.3: Frontier Coding with Emergent Cyber CapabilitiesObserved Checked |
Claude Fable 5 w/ fallback as published by Z.AIClaude Fable 5Exact identityclaude-fable-5-max Canonical product: claude-fable-5 | 83.8% | Relative comparisonprovider-reportedVersion & systemrolling Z.AI GLM-5.3 official comparison table | GLM-5.3: Frontier Coding with Emergent Cyber CapabilitiesObserved Checked |
GPT-5.6 Sol as published by Z.AIGPT-5.6 SolExact identitygpt-5-6-sol-max Canonical product: gpt-5-6-sol | 83.6% | Relative comparisonprovider-reportedVersion & systemrolling Z.AI GLM-5.3 official comparison table | GLM-5.3: Frontier Coding with Emergent Cyber CapabilitiesObserved Checked |
DeepSeek V4 Pro 0813 as published by Z.AIDeepSeek V4 Pro 0813Exact identitydeepseek-v4-pro-0813-max Canonical product: deepseek-v4-pro-0813 | 83.3% | Relative comparisonprovider-reportedVersion & systemrolling Z.AI GLM-5.3 official comparison table | GLM-5.3: Frontier Coding with Emergent Cyber CapabilitiesObserved Checked |
Kimi K3 as published by Z.AIKimi K3Exact identitykimi-k3-max Canonical product: kimi-k3 | 80% | Relative comparisonprovider-reportedVersion & systemrolling Z.AI GLM-5.3 official comparison table | GLM-5.3: Frontier Coding with Emergent Cyber CapabilitiesObserved Checked |
Qwen3.8 Max as published by Z.AIQwen3.8 MaxExact identityqwen-3-8-max-xhigh Canonical product: qwen-3-8-max | 78.5% | Relative comparisonprovider-reportedVersion & systemrolling Z.AI GLM-5.3 official comparison table | GLM-5.3: Frontier Coding with Emergent Cyber CapabilitiesObserved Checked |
Claude Opus 4.8 as published by Z.AIClaude Opus 4.8Exact identityclaude-opus-4-8-max Canonical product: claude-opus-4-8 | 78.1% | Relative comparisonprovider-reportedVersion & systemrolling Z.AI GLM-5.3 official comparison table | GLM-5.3: Frontier Coding with Emergent Cyber CapabilitiesObserved Checked |
GLM-5.2 as published by Z.AIGLM-5.2Exact identityglm-5-2-max Canonical product: glm-5-2 | 77.2% | Relative comparisonprovider-reportedVersion & systemrolling Z.AI GLM-5.3 official comparison table | GLM-5.3: Frontier Coding with Emergent Cyber CapabilitiesObserved Checked |
DeepSeek V4 Pro 0813DeepSeek V4 Pro 0813Exact identitydeepseek-v4-pro-0813-deepseek-0813-release-unspecified Canonical product: deepseek-v4-pro-0813 | 83.3% | Public referenceprovider-reportedVersion & systemrolling DeepSeek V4 Pro 0813 official GA comparison table | DeepSeek permanent refresh sourceObserved Checked |
From result to context
Cybersecurity agent benchmark focused on finding and validating software vulnerabilities under constrained tool environments.
The summary shows one best compatible source result per canonical product. Versions, effort settings and execution systems remain attached to the underlying records.
Scores from different versions or harnesses may not be interchangeable. The published source rows preserve those distinctions and their original units. A source result is not automatically an input to a current index.
This catalogue entry preserves available source evidence. Current indices admit only their specifically reviewed tracks and configurations.
Contamination risk: Low. Lifecycle: Active.