Leaderboard bars
1 ranked models · higher is better
View accessible chart data
| Model | Rank | Provider | Score |
|---|---|---|---|
| Kimi K3 | #1 | Moonshot AI | 63.3% |
Agents · Benchmark profile
Document and spreadsheet workplace QA benchmark measuring office-style multi-file reasoning and calculation.
Data verified 15 July 2026 · Methodology 1.2.0
Visual analysis
Three views of the same sourced leaderboard: model placement, the score curve, and descriptive provider averages. These visuals never alter ranking calculations.
1 ranked models · higher is better
| Model | Rank | Provider | Score |
|---|---|---|---|
| Kimi K3 | #1 | Moonshot AI | 63.3% |
Shows how quickly scores separate down the field
| Model | Rank | Provider | Score |
|---|---|---|---|
| Kimi K3 | #1 | Moonshot AI | 63.3% |
Average within the visible cohort · descriptive only
| Provider | Models | Average score |
|---|---|---|
| Moonshot AI | 1 | 63.3% |
One best score per model · higher is better
| Rank | Model | Provider | License | Score |
|---|---|---|---|---|
| #1 | Kimi K3 kimi-k3 | Moonshot AI | closed | 63.3% |
About OfficeQA Pro
Document and spreadsheet workplace QA benchmark measuring office-style multi-file reasoning and calculation. Results stay tied to the exact model variant and evaluation system. Multiple systems for the same model use the best published score on this page; overall Lumina scoring uses the median of ranking-eligible rows.
Open benchmark source ↗FAQ
Document and spreadsheet workplace QA benchmark measuring office-style multi-file reasoning and calculation.
Kimi K3 by Moonshot AI currently leads with 63.3%.
1 model in the LuminaBench cohort have a qualifying score on this benchmark.
No. This benchmark is display-only and does not enter the overall Lumina composite.
Related