Lumina Capabilities Index
A cross-domain view of model ability, grounded in qualified original benchmark evidence.
Lumina data / Benchmarks
Explore the benchmarks that measure model ability, and the indices that bring the evidence together.
A cross-domain view of model ability, grounded in qualified original benchmark evidence.
Software engineering and programming performance under reviewed task and execution conditions.
Mathematical, scientific and complex problem solving across qualified reasoning evidence.
Tool-driven workflows, terminal execution and professional-task evidence.
Factual reliability, instruction following and document understanding under defined protocols.
Verified open-weight releases, with their original Capabilities scores and research ranges.
The Diamond subset of a graduate-level, expert-written question-answering benchmark.
A broad expert-level academic benchmark covering difficult questions across many domains.
A benchmark for generating scientific-computing code from expert-authored specifications.
Artificial Analysis long-context reasoning benchmark measuring extraction and synthesis over documents from roughly 10k to 100k tokens.
Lumina Verified identifies Lumina’s published combined indices, including the verified Open Weights subset. Provider means retained results reported by a model provider. Independent means an external evaluator or benchmark owner ran the evaluation. A benchmark can contain both kinds of result.
These labels describe provenance, not scoring eligibility or independent replication by Lumina. Unqualified and mirrored records never create an Independent badge. Counts deduplicate model identities; results may span different protocols, configurations and dates.