Score & research range
The point describes the estimate. A retained range describes its documented uncertainty. Some methods do not provide a range.
Research / Methodology
How original measurements become Lumina’s indices. Explore the methods, the benchmarks and what each score can tell you.
01 / The big picture
Original results from all four areas inform a shared capability estimate. The model learns their relationships from the retained evidence.
81 qualified models from the parent index, filtered by verified weight access.
Each specialist index also has its own method. Capabilities uses original evidence across the areas, rather than averaging the four published specialist scores.
02 / Inside each area
Choose an area. Open a benchmark to see what it measures.
Mathematical, scientific and complex problem solving across qualified reasoning evidence.
Graduate-level scientific reasoning under one retained evaluator protocol.
Broad expert reasoning; text-only 2,158-question track, kept separate from tool-enabled results.
Mathematical problem solving on version 2.0.0, private tiers 1–3.
Uses the selected reasoning baseline and qualified early routes. Hollow markers identify provisional points; observed native benchmark results remain separate from predicted components.
Conditional working-model ranges and enabled route-specific prediction ranges retain their actual targets.
Software engineering and programming performance under reviewed task and execution conditions.
Engineering work in a documented client configuration.
Model-plus-harness engineering policy; rubric scores are not pass rates.
Version 1.1 engineering attempts under the retained mini-swe harness.
Scientific programming, background-enabled pass@1 over the retained task set.
One programming task from the qualified LiveBench block; other task children are not independent benchmark definitions.
Combines qualified engineering and programming evidence. Provisional points retain their predicted-component masks; a model-plus-harness result is not an isolated bare-model measurement.
No qualified broad range is available for retained Coding points. A missing range does not remove an otherwise qualified model.
Factual reliability, instruction following and document understanding under defined protocols.
Equal thirds within the General baseline. This is separate from Capabilities.
The observed signed factual-utility component: correct minus incorrect answers, divided by all questions. Not simple accuracy.
The retained instruction quartet mean, not four independent benchmarks.
The observed text-document answer component under LCR v1.1.
An equal-third native composite of factual utility, instruction following and text-document understanding. Signed factual utility rewards correct answers and penalizes incorrect ones; it is not simple accuracy. Supplemental source inputs remain non-scoring.
Complete baselines have no qualified broad interval. Qualified GI prediction ranges remain visible for their specific missing-component targets.
Tool-driven workflows, terminal execution and professional-task evidence.
Tool-created professional deliverables. Native Elo values use a table so no truncated bars imply meaningful score ratios.
Private-set strict workflow success, separate from the AA partial-score track.
Workflow tool execution under the retained April 2026 protocol.
Verified default task executions, separate from alternate pass-count metrics.
Owner terminal-execution track, kept separate from AA-hosted executions.
Enterprise workflow evidence under the retained evaluator protocol.
Default preserves the qualified richer-Workflow population. Professional deliverables, workflow execution and terminal evidence retain their source and harness qualifications.
Qualified conditional ranges are available in the richer-Workflow scope. AG-X retains its narrower missing-X target; AA-only and two-Z source-limited records remain in Research.
03 / Read the result
The point describes the estimate. A retained range describes its documented uncertainty. Some methods do not provide a range.
Established and qualified provisional estimates can appear in the current index. A higher point alone does not establish superiority.
Harnesses, tools and settings stay attached to their measurements. Names and shared branding never make two checkpoints equivalent.
Follow the evidence
Source labels make the benchmark catalogue easier to navigate.
Indices combined and reviewed by Lumina, plus their verified Open Weights subset.
Results reported by a model provider, with their source conditions retained.
Evaluations run by an external evaluator or benchmark owner.