A task with its own rules.
A cybersecurity benchmark that measures whether an agent can turn raw threat-intelligence reports into working detection rules.
- Organisation
- Sakana AI
- Version
- 2026
Agents / Benchmark profile
A cybersecurity benchmark that measures whether an agent can turn raw threat-intelligence reports into working detection rules.
Observed results
1 ranked models · higher is better · labels show rank and score
| Model | Rank | Model ID | Provider | Score | Unit |
|---|---|---|---|---|---|
| Fugu Cyber | #1 | sakana-fugu-cyber | Sakana AI | 72.1 | percent |
One best compatible score per canonical product · higher is better
| Rank | Model | Weights | Evidence | Score |
|---|---|---|---|---|
| 01 | Fugu CyberSakana AI | Closed weights | Source carrier | 72.1% |
Each distinct published measurement is retained. Same-snapshot canonical and configuration projections appear once.
| Published model / configuration | Score | Evidence & protocol | Source & dates |
|---|---|---|---|
Exact BenchLM registry variant Fugu Cyber; bulk export does not retain a complete upstream harness configuration.Fugu CyberExact identitySource label without a registered configuration ID Canonical product: sakana-fugu-cyber | 72.1% | Source carriersource-checkedVersion & system2026 Source-native system | BenchLM public datasets — 2026-08-01Observed Checked |
Exact BenchLM registry variant Fugu Cyber; bulk export does not retain a complete upstream harness configuration.Fugu CyberExact identitySource label without a registered configuration ID Canonical product: sakana-fugu-cyber | 72.1% | Source carriersource-checkedVersion & system2026 Source-native system | BenchLM public datasets — 2026-07-27Observed Checked |
Exact BenchLM registry variant Fugu Cyber; bulk export does not retain a complete upstream harness configuration.Fugu CyberExact identitySource label without a registered configuration ID Canonical product: sakana-fugu-cyber | 72.1% | Source carriersource-checkedVersion & system2026 Source-native system | BenchLM public datasets — 21 July 2026Observed Checked |
From result to context
A cybersecurity benchmark that measures whether an agent can turn raw threat-intelligence reports into working detection rules.
The summary shows one best compatible source result per canonical product. Versions, effort settings and execution systems remain attached to the underlying records.
Scores from different versions or harnesses may not be interchangeable. The published source rows preserve those distinctions and their original units. A source result is not automatically an input to a current index.
This catalogue entry preserves available source evidence. Current indices admit only their specifically reviewed tracks and configurations.
Contamination risk: Unknown. Lifecycle: Active.