A task with its own rules.
A benchmark for browsing agents that must locate difficult-to-find information on the web.
- Organisation
- OpenAI
- Version
- Current / rolling
Research / Benchmark profile
A benchmark for browsing agents that must locate difficult-to-find information on the web.
Observed results
20 ranked models · higher is better · labels show rank and score
| Model | Rank | Model ID | Provider | Score | Unit |
|---|---|---|---|---|---|
| GPT-5.6 Sol | #1 | gpt-5-6-sol | OpenAI | 92.2 | percent |
| Kimi K3 | #2 | kimi-k3 | Moonshot AI | 91.2 | percent |
| Claude Opus 5 | #3 | claude-opus-5 | Anthropic | 90.8 | percent |
| GPT-5.5 | #4 | gpt-5-5 | OpenAI | 90.1 | percent |
| GPT-5.4 | #5 | gpt-5-4 | OpenAI | 89.3 | percent |
| Claude Mythos 5 | #6 | claude-mythos-5 | Anthropic | 88 | percent |
| GPT-5.6 Terra | #7 | gpt-5-6-terra | OpenAI | 87.5 | percent |
| Claude Sonnet 5 | #8 | claude-sonnet-5 | Anthropic | 84.7 | percent |
| Claude Opus 4.8 | #9 | claude-opus-4-8 | Anthropic | 84.3 | percent |
| Claude Opus 4.6 | #10 | claude-opus-4-6 | Anthropic | 83.7 | percent |
| MiniMax M3 | #11 | minimax-m3 | MiniMax | 83.52 | percent |
| DeepSeek V4 Pro | #12 | deepseek-v4-pro | DeepSeek | 83.4 | percent |
| GPT-5.6 Luna | #13 | gpt-5-6-luna | OpenAI | 83.3 | percent |
| Kimi K2.6 | #14 | kimi-k2-6 | Moonshot AI | 83.2 | percent |
| LongCat-2.0 | #15 | longcat-2-0 | Meituan | 79.9 | percent |
| Claude Opus 4.7 | #16 | claude-opus-4-7 | Anthropic | 79.3 | percent |
| Inkling-Small | #17 | inkling-small | Thinking Machines Lab | 77.4 | percent |
| Inkling | #18 | inkling | Thinking Machines Lab | 77.1 | percent |
| Step 3.7 Flash | #19 | step-3-7-flash | StepFun | 75.82 | percent |
| Agents-A1 | #20 | agents-a1 | InternScience | 75.51 | percent |
| DeepSeek V4 Flash | #21 | deepseek-v4-flash | DeepSeek | 73.2 | percent |
| GLM-5.1 | #22 | glm-5-1 | Z.AI | 68 | percent |
| GPT-5.2 | #23 | gpt-5-2 | OpenAI | 65.8 | percent |
| Qwen3.5 122B | #24 | qwen3-5-122b | Alibaba Cloud | 63.8 | percent |
| Qwen3.5-122B-A10B | #25 | qwen3-5-122b-a10b | Alibaba Cloud | 63.8 | percent |
| Qwen3.5 397B A17B | #26 | qwen3-5-397b | Alibaba Cloud | 62 | percent |
| Qwen3.5 27B | #27 | qwen3-5-27b | Alibaba Cloud | 61 | percent |
| Qwen3.5 35B | #28 | qwen3-5-35b | Alibaba Cloud | 61 | percent |
| Qwen3.5-35B-A3B | #29 | qwen3-5-35b-a3b | Alibaba Cloud | 61 | percent |
| Kimi K2.5 | #30 | kimi-k2-5 | Moonshot AI | 60.6 | percent |
| GLM-4.7 | #31 | glm-4-7 | Z.AI | 52 | percent |
| Nemotron 3 Ultra | #32 | nemotron-3-ultra | NVIDIA | 44.4 | percent |
| NVIDIA Nemotron 3.5 Lightning 30B-A3B | #33 | nemotron-3-5-lightning-30b-a3b | NVIDIA | 36.97 | percent |
One best compatible score per canonical product · higher is better
| Rank | Model | Weights | Evidence | Score |
|---|---|---|---|---|
| 01 | GPT-5.6 SolOpenAI | Closed weights | Source carrier | 92.2% |
| 02 | Kimi K3Moonshot AI | Open weights | Direct source result | 91.2% |
| 03 | Claude Opus 5Anthropic | Closed weights | Direct source result | 90.8% |
| 04 | GPT-5.5OpenAI | Closed weights | Source carrier | 90.1% |
| 05 | GPT-5.4OpenAI | Closed weights | Source carrier | 89.3% |
| 06 | Claude Mythos 5Anthropic | Closed weights | Source carrier | 88% |
| 07 | GPT-5.6 TerraOpenAI | Closed weights | Source carrier | 87.5% |
| 08 | Claude Sonnet 5Anthropic | Closed weights | Source carrier | 84.7% |
| 09 | Claude Opus 4.8Anthropic | Closed weights | Source carrier | 84.3% |
| 10 | Claude Opus 4.6Anthropic | Closed weights | Source carrier | 83.7% |
Each distinct published measurement is retained. Same-snapshot canonical and configuration projections appear once.
| Published model / configuration | Score | Evidence & protocol | Source & dates |
|---|---|---|---|
NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16; NVIDIA release evaluation; temperature=1.0; top_p=0.95NVIDIA Nemotron 3.5 Lightning 30B-A3BExact identitynemotron-3-5-lightning-30b-a3b-default Canonical product: nemotron-3-5-lightning-30b-a3b | 37.0% | Direct source resultprovider-reportedVersion & systemCurrent NVIDIA NeMo Gym / NeMo Evaluator SDK consistent release harness | NVIDIA Nemotron 3.5 Lightning 30B-A3B BF16 model cardObserved Checked |
Exact BenchLM registry variant GPT-5.6 Sol; bulk export does not retain a complete upstream harness configuration.GPT-5.6 SolExact identitySource label without a registered configuration ID Canonical product: gpt-5-6-sol | 92.2% | Source carriersource-checkedVersion & system2025 Source-native system | BenchLM public datasets — 2026-08-01Observed Checked |
Exact BenchLM registry variant Kimi K3; bulk export does not retain a complete upstream harness configuration.Kimi K3Exact identitySource label without a registered configuration ID Canonical product: kimi-k3 | 91.2% | Source carriersource-checkedVersion & system2025 Source-native system | BenchLM public datasets — 2026-08-01Observed Checked |
Exact BenchLM registry variant Claude Opus 5; bulk export does not retain a complete upstream harness configuration.Claude Opus 5Exact identitySource label without a registered configuration ID Canonical product: claude-opus-5 | 90.8% | Source carriersource-checkedVersion & system2025 Source-native system | BenchLM public datasets — 2026-08-01Observed Checked |
Exact BenchLM registry variant GPT-5.5 Pro; bulk export does not retain a complete upstream harness configuration.GPT-5.5Exact identitygpt-5-5-pro-default-high Canonical product: gpt-5-5 | 90.1% | Source carriersource-checkedVersion & system2025 Source-native system | BenchLM public datasets — 2026-08-01Observed Checked |
Exact BenchLM registry variant GPT-5.4 Pro; bulk export does not retain a complete upstream harness configuration.GPT-5.4Exact identitygpt-5-4-pro-default-medium Canonical product: gpt-5-4 | 89.3% | Source carriersource-checkedVersion & system2025 Source-native system | BenchLM public datasets — 2026-08-01Observed Checked |
Exact BenchLM registry variant Claude Mythos 5; bulk export does not retain a complete upstream harness configuration.Claude Mythos 5Exact identitySource label without a registered configuration ID Canonical product: claude-mythos-5 | 88% | Source carriersource-checkedVersion & system2025 Source-native system | BenchLM public datasets — 2026-08-01Observed Checked |
Exact BenchLM registry variant GPT-5.6 Terra; bulk export does not retain a complete upstream harness configuration.GPT-5.6 TerraExact identitySource label without a registered configuration ID Canonical product: gpt-5-6-terra | 87.5% | Source carriersource-checkedVersion & system2025 Source-native system | BenchLM public datasets — 2026-08-01Observed Checked |
Exact BenchLM registry variant Claude Sonnet 5; bulk export does not retain a complete upstream harness configuration.Claude Sonnet 5Exact identitySource label without a registered configuration ID Canonical product: claude-sonnet-5 | 84.7% | Source carriersource-checkedVersion & system2025 Source-native system | BenchLM public datasets — 2026-08-01Observed Checked |
Exact BenchLM registry variant GPT-5.5; bulk export does not retain a complete upstream harness configuration.GPT-5.5Exact identitySource label without a registered configuration ID Canonical product: gpt-5-5 | 84.4% | Source carriersource-checkedVersion & system2025 Source-native system | BenchLM public datasets — 2026-08-01Observed Checked |
From result to context
A benchmark for browsing agents that must locate difficult-to-find information on the web.
The summary shows one best compatible source result per canonical product. Versions, effort settings and execution systems remain attached to the underlying records.
Scores from different versions or harnesses may not be interchangeable. The published source rows preserve those distinctions and their original units. A source result is not automatically an input to a current index.
This catalogue entry preserves available source evidence. Current indices admit only their specifically reviewed tracks and configurations.
Contamination risk: Low. Lifecycle: Active.