LuminaXspace import (16–21 July originals)
Imported 24 additional original @LuminaXspace posts as editorial-only News records (exact text, media, IDs). Replies excluded. News remains structurally excluded from scoring. Registry now has 66 articles.
Append-only release history
Material source, methodology, ranking, correction and product events remain visible. Historical snapshots are appended rather than overwritten.
Imported 24 additional original @LuminaXspace posts as editorial-only News records (exact text, media, IDs). Replies excluded. News remains structurally excluded from scoring. Registry now has 66 articles.
Appended source-checked BenchLM public multi-model rows for exact registry variants: densified BrowseComp and MMLU-Pro sparse cells, filled missing Terminal-Bench and GDPval-AA entries. No invented scores. Snapshot: 103 models, 1987 raw results, 84 families, 177 sources.
Admitted Moonshot Kimi K3 with source-backed pricing, Artificial Analysis speed/TTFT, and ranking-eligible results across Terminal-Bench, BrowseComp, MCP Atlas, Toolathlon, GDPval-AA, HLE, GPQA, MMMU-Pro, CharXiv, SciCode, AA-LCR, CritPt, Omniscience, and related agent families. Registered six new coding reference families. Listed Alibaba Qwen3.8 Max Preview without invented scores. Snapshot ranks 102 of 103 models.
Regenerated methodology 1.4.1 Power Index: Kimi K3 enters the top tier (rank #3 in this snapshot). Models with zero ranking-eligible evidence and no consensus prior stay Not evaluated instead of receiving a neutral-floor score.
Leaderboard and model directory provider filters are now horizontal lab chips with each lab’s brand colour, keyboard radiogroup navigation, and URL-synced selection.
Missing core families are now imputed from consensus priors (discounted) instead of pure weight renormalization. Stricter reliability, weaker Cursor substitutes, real-Cursor finishing bonus only, and a critical-family coverage penalty. Fixes GPT-5.2 above GPT-5.5 and Qwen3.7 Max above Grok 4.5 / GLM-5.2. Efficiency and price still never affect Power.
Overall leaderboard rank is pure Power from a CursorBench-heavy core suite (SWE, Terminal, HLE, GPQA, GDPval, and related). Superseded by 1.4.1 imputation rules for missing cores. Efficiency and list price never affect Power.
Added 21 further source-checked variants (Grok 4 family, Kimi K2 Thinking, Qwen3 Max/Omni/35B, MiniMax M2.x, Claude Sonnet 4/3.7, GLM-4.6, Gemini 2.5, o4-mini/o1, Nemotron 3 Super, Ling 2.6, and more). Registry is now 103 models with 1987 raw results across 84 benchmark families. Decision cohort target reached; further growth is optional refinement only.
Added 29 source-checked variants from Artificial Analysis independent evaluation payloads (OpenAI, Anthropic, Moonshot, xAI, Alibaba, DeepSeek, Z.AI, MiniMax, Xiaomi, NVIDIA, Google, Cohere, Amazon). Product policy targets ~100 ranked models, not a full 300+ catalog.
Mirrored additional public BenchLM leaderboard rows only. Registry now has 84 families and 1987 raw results across 103 models. Empty families remain empty where no public cohort rows exist.
Expanded to 103 models. Overall rank blends direct benchmarks, BenchLM category priors for sparse cells, sample-size damping, and CursorBench coding-agent signal. Frontier order targets Mythos → Fable → Sol → Opus → Grok. Snapshot ranks 102 models from 1987 raw results.
Added Meta Muse Spark, Xiaomi MiMo, Z.AI GLM-5 family, Gemini 3 Pro, Composer-class peers and more from the BenchLM July snapshot, plus CursorBench as a weighted coding-agent family.
Overall Lumina rank became capability-only. Superseded by 1.3.0 consensus blend and CursorBench calibration.
Densified public BenchLM leaderboard rows across the cohort: 1987 raw results spanning the majority of the registry. Model profiles show best-per-benchmark score tables linked into full per-benchmark leaderboards.
Mirrored public BenchLM benchmark leaderboard rows for the exact model cohort, densifying HLE, GPQA, Terminal-Bench, SWE-Bench, OSWorld, BrowseComp and related families.
Ordinal ranks applied to every model with a computable Lumina score. Incomplete dimensions renormalized; superseded by 1.2.0 for overall composite definition.
Initial source-qualified snapshot under the complete-dimension official ranking gate. Superseded by methodology 1.1.0 for publication status only; raw evidence rows are preserved.
Removed the former X-derived observation store and every public reference. LuminaXspace posts remain exclusively in News and can never enter benchmark evidence or scoring.
Fixed the 70/15/15 Lumina composite, six capability-category weights, median multi-system aggregation, coverage and confidence rules, provisional scoring, and the five-family/three-category official ranking gate.
Added source-qualified model, price, benchmark-version, evaluation-system, result and operational records with append-only manifests and review reports.
Added 102 current model variants. Official specifications and comparable list prices were stored with check dates; missing fields remained unavailable.
Registered the initial benchmark families with official project links before scoring eligibility was activated.
Imported 66 recent original posts as editorial records. News remains structurally excluded from scoring.
Created the initial provider source records; the current registry now contains 177 source records and 84 benchmark families.