A task with its own rules.
Multi-needle long-context retrieval and reasoning benchmark stressing multi-document evidence gathering.
- Organisation
- OpenAI
- Version
- 2
Research / Benchmark profile
Multi-needle long-context retrieval and reasoning benchmark stressing multi-document evidence gathering.
Observed results
10 ranked models · higher is better · labels show rank and score
| Model | Rank | Model ID | Provider | Score | Unit |
|---|---|---|---|---|---|
| Sakana Fugu-Ultra | #1 | sakana-fugu-ultra | Sakana AI | 93.6 | percent |
| Gemini 3.6 Flash | #2 | gemini-3-6-flash | 91.8 | percent | |
| Qwen3.7-Plus | #3 | qwen-3-7-plus | Alibaba Cloud | 91.7 | percent |
| Qwen3.7-Max | #4 | qwen-3-7-max | Alibaba Cloud | 90.4 | percent |
| Sakana Fugu | #5 | sakana-fugu | Sakana AI | 86.6 | percent |
| Gemini 3.5 Flash | #6 | gemini-3-5-flash | 77.3 | percent | |
| Gemini 3.5 Flash-Lite | #7 | gemini-3-5-flash-lite | 72.2 | percent | |
| Gemini 3.1 Flash-Lite | #8 | gemini-3-1-flash-lite | 60.1 | percent | |
| Muse Spark 1.1 | #9 | muse-spark-1-1 | Meta | 54.1 | percent |
| Gemma 4 12B Unified | #10 | gemma-4-12b | 43.4 | percent |
One best compatible score per canonical product · higher is better
| Rank | Model | Weights | Evidence | Score |
|---|---|---|---|---|
| 01 | Sakana Fugu-UltraSakana AI | Closed weights | Source carrier | 93.6% |
| 02 | Gemini 3.6 FlashGoogle | Closed weights | Public reference | 91.8% |
| 03 | Qwen3.7-PlusAlibaba Cloud | Closed weights | Source carrier | 91.7% |
| 04 | Qwen3.7-MaxAlibaba Cloud | Closed weights | Source carrier | 90.4% |
| 05 | Sakana FuguSakana AI | Closed weights | Source carrier | 86.6% |
| 06 | Gemini 3.5 FlashGoogle | Closed weights | Source carrier | 77.3% |
| 07 | Gemini 3.5 Flash-LiteGoogle | Closed weights | Public reference | 72.2% |
| 08 | Gemini 3.1 Flash-LiteGoogle | Closed weights | Public reference | 60.1% |
| 09 | Muse Spark 1.1Meta | Closed weights | Public reference | 54.1% |
| 10 | Gemma 4 12B UnifiedGoogle | Open weights | Source carrier | 43.4% |
Each distinct published measurement is retained. Same-snapshot canonical and configuration projections appear once.
| Published model / configuration | Score | Evidence & protocol | Source & dates |
|---|---|---|---|
Exact BenchLM registry variant Sakana Fugu-Ultra; bulk export does not retain a complete upstream harness configuration.Sakana Fugu-UltraExact identitySource label without a registered configuration ID Canonical product: sakana-fugu-ultra | 93.6% | Source carriersource-checkedVersion & system2025 Source-native system | BenchLM public datasets — 2026-08-01Observed Checked |
Exact BenchLM registry variant Qwen3.7 Plus; bulk export does not retain a complete upstream harness configuration.Qwen3.7-PlusExact identitySource label without a registered configuration ID Canonical product: qwen-3-7-plus | 91.7% | Source carriersource-checkedVersion & system2025 Source-native system | BenchLM public datasets — 2026-08-01Observed Checked |
Exact BenchLM registry variant Qwen3.7 Max; bulk export does not retain a complete upstream harness configuration.Qwen3.7-MaxExact identitySource label without a registered configuration ID Canonical product: qwen-3-7-max | 90.4% | Source carriersource-checkedVersion & system2025 Source-native system | BenchLM public datasets — 2026-08-01Observed Checked |
Exact BenchLM registry variant Sakana Fugu; bulk export does not retain a complete upstream harness configuration.Sakana FuguExact identitySource label without a registered configuration ID Canonical product: sakana-fugu | 86.6% | Source carriersource-checkedVersion & system2025 Source-native system | BenchLM public datasets — 2026-08-01Observed Checked |
Exact BenchLM registry variant Gemini 3.5 Flash; bulk export does not retain a complete upstream harness configuration.Gemini 3.5 FlashExact identitySource label without a registered configuration ID Canonical product: gemini-3-5-flash | 77.3% | Source carriersource-checkedVersion & system2025 Source-native system | BenchLM public datasets — 2026-08-01Observed Checked |
Exact BenchLM registry variant Gemini 3.5 Flash-Lite; bulk export does not retain a complete upstream harness configuration.Gemini 3.5 Flash-LiteExact identitySource label without a registered configuration ID Canonical product: gemini-3-5-flash-lite | 72.2% | Source carriersource-checkedVersion & system2025 Source-native system | BenchLM public datasets — 2026-08-01Observed Checked |
Exact BenchLM registry variant Gemma 4 12B; bulk export does not retain a complete upstream harness configuration.Gemma 4 12B UnifiedExact identitySource label without a registered configuration ID Canonical product: gemma-4-12b | 43.4% | Source carriersource-checkedVersion & system2025 Source-native system | BenchLM public datasets — 2026-08-01Observed Checked |
Exact BenchLM registry variant Sakana Fugu-Ultra; bulk export does not retain a complete upstream harness configuration.Sakana Fugu-UltraExact identitySource label without a registered configuration ID Canonical product: sakana-fugu-ultra | 93.6% | Source carriersource-checkedVersion & system2025 Source-native system | BenchLM public datasets — 2026-07-27Observed Checked |
Exact BenchLM registry variant Qwen3.7 Plus; bulk export does not retain a complete upstream harness configuration.Qwen3.7-PlusExact identitySource label without a registered configuration ID Canonical product: qwen-3-7-plus | 91.7% | Source carriersource-checkedVersion & system2025 Source-native system | BenchLM public datasets — 2026-07-27Observed Checked |
Exact BenchLM registry variant Qwen3.7 Max; bulk export does not retain a complete upstream harness configuration.Qwen3.7-MaxExact identitySource label without a registered configuration ID Canonical product: qwen-3-7-max | 90.4% | Source carriersource-checkedVersion & system2025 Source-native system | BenchLM public datasets — 2026-07-27Observed Checked |
From result to context
Multi-needle long-context retrieval and reasoning benchmark stressing multi-document evidence gathering.
The summary shows one best compatible source result per canonical product. Versions, effort settings and execution systems remain attached to the underlying records.
Scores from different versions or harnesses may not be interchangeable. The published source rows preserve those distinctions and their original units. A source result is not automatically an input to a current index.
This catalogue entry preserves available source evidence. Current indices admit only their specifically reviewed tracks and configurations.
Contamination risk: Low. Lifecycle: Active.