A task with its own rules.
A realistic cybersecurity benchmark for evaluating whether agents can turn known software vulnerabilities into concrete attacks.
- Organisation
- ExploitGym authors
- Version
- 2026-05
Agents / Benchmark profile
A realistic cybersecurity benchmark for evaluating whether agents can turn known software vulnerabilities into concrete attacks.
Observed results
8 ranked models · higher is better · labels show rank and score
| Model | Rank | Model ID | Provider | Score | Unit |
|---|---|---|---|---|---|
| GLM-5.3 | #1 | glm-5-3 | Z.AI | 130 | tasks |
| GPT-5.6 Sol | #2 | gpt-5-6-sol | OpenAI | 33.7 | percent |
| GPT-5.6 Terra | #3 | gpt-5-6-terra | OpenAI | 23.2 | percent |
| Claude Mythos 5 | #4 | claude-mythos-5 | Anthropic | 17.5 | percent |
| GPT-5.5 | #5 | gpt-5-5 | OpenAI | 13.4 | percent |
| GPT-5.6 Luna | #6 | gpt-5-6-luna | OpenAI | 12.4 | percent |
| GPT-5.4 | #7 | gpt-5-4 | OpenAI | 6 | percent |
| Muse Spark 1.1 | #8 | muse-spark-1-1 | Meta | 0.8 | percent |
One best compatible score per canonical product · higher is better
| Rank | Model | Weights | Evidence | Score |
|---|---|---|---|---|
| 01 | GLM-5.3Z.AI | Open announced | Public reference | 130 tasks |
| 02 | GPT-5.6 SolOpenAI | Closed weights | Source carrier | 33.7% |
| 03 | GPT-5.6 TerraOpenAI | Closed weights | Source carrier | 23.2% |
| 04 | Claude Mythos 5Anthropic | Closed weights | Source carrier | 17.5% |
| 05 | GPT-5.5OpenAI | Closed weights | Source carrier | 13.4% |
| 06 | GPT-5.6 LunaOpenAI | Closed weights | Source carrier | 12.4% |
| 07 | GPT-5.4OpenAI | Closed weights | Source carrier | 6% |
| 08 | Muse Spark 1.1Meta | Closed weights | Public reference | 0.8% |
Each distinct published measurement is retained. Same-snapshot canonical and configuration projections appear once.
| Published model / configuration | Score | Evidence & protocol | Source & dates |
|---|---|---|---|
GPT-5.6 Sol as published by Z.AIGPT-5.6 SolExact identitygpt-5-6-sol-max Canonical product: gpt-5-6-sol | 293 tasks | Public referenceprovider-reportedVersion & system6h Z.AI GLM-5.3 official comparison table | GLM-5.3: Frontier Coding with Emergent Cyber CapabilitiesObserved Checked |
Claude Fable 5 w/ fallback as published by Z.AIClaude Fable 5Exact identityclaude-fable-5-max Canonical product: claude-fable-5 | 247 tasks | Public referenceprovider-reportedVersion & system6h Z.AI GLM-5.3 official comparison table | GLM-5.3: Frontier Coding with Emergent Cyber CapabilitiesObserved Checked |
GPT-5.6 Sol as published by Z.AIGPT-5.6 SolExact identitygpt-5-6-sol-max Canonical product: gpt-5-6-sol | 216 tasks | Public referenceprovider-reportedVersion & system2h Z.AI GLM-5.3 official comparison table | GLM-5.3: Frontier Coding with Emergent Cyber CapabilitiesObserved Checked |
Claude Fable 5 w/ fallback as published by Z.AIClaude Fable 5Exact identityclaude-fable-5-max Canonical product: claude-fable-5 | 181 tasks | Public referenceprovider-reportedVersion & system2h Z.AI GLM-5.3 official comparison table | GLM-5.3: Frontier Coding with Emergent Cyber CapabilitiesObserved Checked |
GLM-5.3 (max)GLM-5.3Exact identityglm-5-3-max Canonical product: glm-5-3 | 130 tasks | Public referenceprovider-reportedVersion & system6h Z.AI GLM-5.3 official comparison table | GLM-5.3: Frontier Coding with Emergent Cyber CapabilitiesObserved Checked |
Claude Opus 4.8 as published by Z.AIClaude Opus 4.8Exact identityclaude-opus-4-8-max Canonical product: claude-opus-4-8 | 120 tasks | Public referenceprovider-reportedVersion & system6h Z.AI GLM-5.3 official comparison table | GLM-5.3: Frontier Coding with Emergent Cyber CapabilitiesObserved Checked |
GLM-5.3 (max)GLM-5.3Exact identityglm-5-3-max Canonical product: glm-5-3 | 105 tasks | Public referenceprovider-reportedVersion & system2h Z.AI GLM-5.3 official comparison table | GLM-5.3: Frontier Coding with Emergent Cyber CapabilitiesObserved Checked |
Claude Opus 4.8 as published by Z.AIClaude Opus 4.8Exact identityclaude-opus-4-8-max Canonical product: claude-opus-4-8 | 80 tasks | Public referenceprovider-reportedVersion & system2h Z.AI GLM-5.3 official comparison table | GLM-5.3: Frontier Coding with Emergent Cyber CapabilitiesObserved Checked |
Kimi K3 as published by Z.AIKimi K3Exact identitykimi-k3-max Canonical product: kimi-k3 | 70 tasks | Public referenceprovider-reportedVersion & system6h Z.AI GLM-5.3 official comparison table | GLM-5.3: Frontier Coding with Emergent Cyber CapabilitiesObserved Checked |
GLM-5.2 as published by Z.AIGLM-5.2Exact identityglm-5-2-max Canonical product: glm-5-2 | 39 tasks | Public referenceprovider-reportedVersion & system6h Z.AI GLM-5.3 official comparison table | GLM-5.3: Frontier Coding with Emergent Cyber CapabilitiesObserved Checked |
From result to context
A realistic cybersecurity benchmark for evaluating whether agents can turn known software vulnerabilities into concrete attacks.
The summary shows one best compatible source result per canonical product. Versions, effort settings and execution systems remain attached to the underlying records.
Scores from different versions or harnesses may not be interchangeable. The published source rows preserve those distinctions and their original units. A source result is not automatically an input to a current index.
This catalogue entry preserves available source evidence. Current indices admit only their specifically reviewed tracks and configurations.
Contamination risk: Low. Lifecycle: Active.