Leaderboard bars
4 ranked models · higher is better
View accessible chart data
| Model | Rank | Provider | Score |
|---|---|---|---|
| Claude Fable 5 | #1 | Anthropic | 38.6% |
| GPT-5.5 | #2 | OpenAI | 36.2% |
| Gemini 3.5 Flash | #3 | 33.6% | |
| Gemini 3.1 Pro Preview | #4 | 26.5% |
Multimodal · Benchmark profile
Measures agentic spatial reasoning over blueprint-style visual information.
Data verified 15 July 2026 · Methodology 1.2.0
Benchmark score on Blueprint-Bench 2
Claude Fable 5 leads at 38.6%, followed by GPT-5.5 (36.2%) and Gemini 3.5 Flash (33.6%).
Visual analysis
Three views of the same sourced leaderboard: model placement, the score curve, and descriptive provider averages. These visuals never alter ranking calculations.
4 ranked models · higher is better
| Model | Rank | Provider | Score |
|---|---|---|---|
| Claude Fable 5 | #1 | Anthropic | 38.6% |
| GPT-5.5 | #2 | OpenAI | 36.2% |
| Gemini 3.5 Flash | #3 | 33.6% | |
| Gemini 3.1 Pro Preview | #4 | 26.5% |
Shows how quickly scores separate down the field
| Model | Rank | Provider | Score |
|---|---|---|---|
| Claude Fable 5 | #1 | Anthropic | 38.6% |
| GPT-5.5 | #2 | OpenAI | 36.2% |
| Gemini 3.5 Flash | #3 | 33.6% | |
| Gemini 3.1 Pro Preview | #4 | 26.5% |
Average within the visible cohort · descriptive only
| Provider | Models | Average score |
|---|---|---|
| Anthropic | 1 | 38.6% |
| OpenAI | 1 | 36.2% |
| 2 | 30.1% |
One best score per model · higher is better
| Rank | Model | Provider | License | Score |
|---|---|---|---|---|
| #1 | Claude Fable 5 claude-fable-5 | Anthropic | closed | 38.6% |
| #2 | GPT-5.5 gpt-5.5 | OpenAI | closed | 36.2% |
| #3 | Gemini 3.5 Flash gemini-3.5-flash | closed | 33.6% | |
| #4 | Gemini 3.1 Pro Preview gemini-3.1-pro-preview | closed | 26.5% |
The top of this snapshot is led by Claude Fable 5 at 38.6%; third place is 5.0 points behind. The top-4 spread is 12.1 points.
About Blueprint-Bench 2
Measures agentic spatial reasoning over blueprint-style visual information. Results stay tied to the exact model variant and evaluation system. Multiple systems for the same model use the best published score on this page; overall Lumina scoring uses the median of ranking-eligible rows.
Open benchmark source ↗FAQ
Measures agentic spatial reasoning over blueprint-style visual information.
Claude Fable 5 by Anthropic currently leads with 38.6%.
4 models in the LuminaBench cohort have a qualifying score on this benchmark.
No. This benchmark is display-only and does not enter the overall Lumina composite.
Related