A task with its own rules.
An agentic browsing benchmark where models search the web, gather evidence, and answer list-style questions using browser tools.
- Organisation
- Meta AI
- Version
- 2026
Agents / Benchmark profile
An agentic browsing benchmark where models search the web, gather evidence, and answer list-style questions using browser tools.
Observed results
12 ranked models · higher is better · labels show rank and score
| Model | Rank | Model ID | Provider | Score | Unit |
|---|---|---|---|---|---|
| Claude Opus 5 | #1 | claude-opus-5 | Anthropic | 95 | percent |
| Kimi K3 | #2 | kimi-k3 | Moonshot AI | 95 | percent |
| Claude Opus 4.8 | #3 | claude-opus-4-8 | Anthropic | 93.1 | percent |
| Step 3.7 Flash | #4 | step-3-7-flash | StepFun | 92.82 | percent |
| Kimi K2.6 | #5 | kimi-k2-6 | Moonshot AI | 92.5 | percent |
| Muse Spark 1.1 | #6 | muse-spark-1-1 | Meta | 84.9 | percent |
| Kimi K2.5 | #7 | kimi-k2-5 | Moonshot AI | 77.1 | percent |
| Muse Spark | #8 | muse-spark | Meta | 74.8 | percent |
| Muse Glimmer 30B | #9 | muse-glimmer-30b | Meta | 74.6 | percent |
| Claude Opus 4.6 | #10 | claude-opus-4-6 | Anthropic | 73.7 | percent |
| GPT-5.4 | #11 | gpt-5-4 | OpenAI | 73.6 | percent |
| Grok 4.20 | #12 | grok-4-20 | xAI | 62.8 | percent |
One best compatible score per canonical product · higher is better
| Rank | Model | Weights | Evidence | Score |
|---|---|---|---|---|
| 01 | Claude Opus 5Anthropic | Closed weights | Source carrier | 95% |
| 02 | Kimi K3Moonshot AI | Open weights | Source carrier | 95% |
| 03 | Claude Opus 4.8Anthropic | Closed weights | Source carrier | 93.1% |
| 04 | Step 3.7 FlashStepFun | Open weights | Source carrier | 92.8% |
| 05 | Kimi K2.6Moonshot AI | Open weights | Source carrier | 92.5% |
| 06 | Muse Spark 1.1Meta | Closed weights | Source carrier | 84.9% |
| 07 | Kimi K2.5Moonshot AI | Open weights | Source carrier | 77.1% |
| 08 | Muse SparkMeta | Closed weights | Source carrier | 74.8% |
| 09 | Muse Glimmer 30BMeta | Open weights | Public reference | 74.6% |
| 10 | Claude Opus 4.6Anthropic | Closed weights | Source carrier | 73.7% |
Each distinct published measurement is retained. Same-snapshot canonical and configuration projections appear once.
| Published model / configuration | Score | Evidence & protocol | Source & dates |
|---|---|---|---|
Muse Glimmer-30B; high reasoning; temperature=1.0; top_p=0.95; top_k=64Muse Glimmer 30BExact identitymuse-glimmer-30b-high Canonical product: muse-glimmer-30b | 74.6% | Public referenceprovider-reportedVersion & system900 questions Meta autonomous browsing evaluation averaged across four runs | Muse Glimmer Evaluation MethodologyObserved Checked |
Exact BenchLM registry variant Claude Opus 5; bulk export does not retain a complete upstream harness configuration.Claude Opus 5Exact identitySource label without a registered configuration ID Canonical product: claude-opus-5 | 95% | Source carriersource-checkedVersion & system2026 Source-native system | BenchLM public datasets — 2026-08-01Observed Checked |
Kimi K3 public default configuration (reasoning_effort=max; temperature=1.0)Kimi K3Exact identitykimi-k3-max Canonical product: kimi-k3 | 95% | Source carriersource-checkedVersion & systemDeepSearchQA F1 Moonshot Kimi K3 model-card evaluation | BenchLM public datasets — 2026-08-01Observed Checked |
Exact BenchLM registry variant Claude Opus 4.8; bulk export does not retain a complete upstream harness configuration.Claude Opus 4.8Exact identitySource label without a registered configuration ID Canonical product: claude-opus-4-8 | 93.1% | Source carriersource-checkedVersion & system2026 Source-native system | BenchLM public datasets — 2026-08-01Observed Checked |
Exact BenchLM registry variant Step 3.7 Flash; bulk export does not retain a complete upstream harness configuration.Step 3.7 FlashExact identitySource label without a registered configuration ID Canonical product: step-3-7-flash | 92.8% | Source carriersource-checkedVersion & system2026 Source-native system | BenchLM public datasets — 2026-08-01Observed Checked |
Exact BenchLM registry variant Kimi K2.6; bulk export does not retain a complete upstream harness configuration.Kimi K2.6Exact identitySource label without a registered configuration ID Canonical product: kimi-k2-6 | 92.5% | Source carriersource-checkedVersion & system2026 Source-native system | BenchLM public datasets — 2026-08-01Observed Checked |
Exact BenchLM registry variant Muse Spark 1.1; bulk export does not retain a complete upstream harness configuration.Muse Spark 1.1Exact identitySource label without a registered configuration ID Canonical product: muse-spark-1-1 | 84.9% | Source carriersource-checkedVersion & system2026 Source-native system | BenchLM public datasets — 2026-08-01Observed Checked |
Exact BenchLM registry variant Kimi K2.5; bulk export does not retain a complete upstream harness configuration.Kimi K2.5Exact identitySource label without a registered configuration ID Canonical product: kimi-k2-5 | 77.1% | Source carriersource-checkedVersion & system2026 Source-native system | BenchLM public datasets — 2026-08-01Observed Checked |
Exact BenchLM registry variant Muse Spark; bulk export does not retain a complete upstream harness configuration.Muse SparkExact identitySource label without a registered configuration ID Canonical product: muse-spark | 74.8% | Source carriersource-checkedVersion & system2026 Source-native system | BenchLM public datasets — 2026-08-01Observed Checked |
Exact BenchLM registry variant Claude Opus 4.6; bulk export does not retain a complete upstream harness configuration.Claude Opus 4.6Exact identitySource label without a registered configuration ID Canonical product: claude-opus-4-6 | 73.7% | Source carriersource-checkedVersion & system2026 Source-native system | BenchLM public datasets — 2026-08-01Observed Checked |
From result to context
An agentic browsing benchmark where models search the web, gather evidence, and answer list-style questions using browser tools.
The summary shows one best compatible source result per canonical product. Versions, effort settings and execution systems remain attached to the underlying records.
Scores from different versions or harnesses may not be interchangeable. The published source rows preserve those distinctions and their original units. A source result is not automatically an input to a current index.
This catalogue entry preserves available source evidence. Current indices admit only their specifically reviewed tracks and configurations.
Contamination risk: Unknown. Lifecycle: Active.