This model│ Other models · 155 ratedScore and research range
Full-precision point
104.54207756168418
Range
100.01–109.07 · 90% Research Range
Scope
R0 latent score conditional on frozen working model and native aggregate portfolio
Score uses the admitted direct native evidence under the selected method.
Coding
No current rating
Software engineering and programming.
Awaiting a qualified category score
General
No current rating
Factual reliability and instruction following.
Awaiting a qualified category score
Agentic
No current rating
Tool use, workflows and professional tasks.
Awaiting a qualified category score
14catalogue benchmarks
15catalogue sources
Evidence availableNo score is inferred from missing evidence.
Current index snapshot · 2026-09-16. Each category uses its own scale; scores across categories are not directly comparable. Counts describe published catalogue evidence. A point lead does not establish superiority.
o4-mini has no current Capabilities rating. Explore the rated models below.
Lumina Capabilities Index
Capabilities versus API price
The current Capabilities cohort, with available operational measurements.
Lumina
2025+ · 62 models
Chart loads as you explore
Lumina Capabilities Index
Capabilities versus output speed
The current Capabilities cohort, with available operational measurements.
Lumina
2025+ · 72 models · Developer API · 10K
Chart loads as you explore
Explore the detail
Sources & model details.
Original evidence, with its context.
Current index evidence 1Scores, research ranges and the sources behind each category.
Reasoning
Score uses the admitted direct native evidence under the selected method.
Exact checkpoint
o4-mini
Method
methods-a4471cee84fd478d2ed3
Range interpretation
R0 latent score conditional on frozen working model and native aggregate portfolio
Sources
aa · epoch
Source cutoff
{"asOfMeaning":"Selected local candidate: retained historical measurements plus reviewed additions; execution dates remain source-native.","newIdentityRecovery":"2026-09-13T08:41:08.804270+00:00","publicProjectionDate":"2026-09-16T06:02:45.491398+00:00","scoring":"Inherited pinned baseline source vintages plus retained4A/4A.1 originals and SWE owner correction; heterogeneous execution dates, often unknown. Not a claim of simultaneous2026-09-13 execution."}
Specialist task evidenceSeparate agent runs and execution conditions.
Specialist evidence
Agent runs and task economics
Source-native operational evidence for o4-mini. Exact configurations remain separate. Costs are comparable only within the same selected benchmark/evaluation, and these rows never enter Overall Score.