Leaderboard bars
2 ranked models · higher is better
View accessible chart data
| Model | Rank | Provider | Score |
|---|---|---|---|
| Kimi K2.5 | #1 | Moonshot AI | 63.5% |
| MiniMax M3 | #2 | MiniMax | 52.6% |
Research · Benchmark profile
An evaluation of agents attempting to replicate machine-learning research from published papers.
Data verified 15 July 2026 · Methodology 1.2.0
Benchmark score on PaperBench
Kimi K2.5 leads at 63.5%, followed by MiniMax M3 (52.6%).
Visual analysis
Three views of the same sourced leaderboard: model placement, the score curve, and descriptive provider averages. These visuals never alter ranking calculations.
2 ranked models · higher is better
| Model | Rank | Provider | Score |
|---|---|---|---|
| Kimi K2.5 | #1 | Moonshot AI | 63.5% |
| MiniMax M3 | #2 | MiniMax | 52.6% |
Shows how quickly scores separate down the field
| Model | Rank | Provider | Score |
|---|---|---|---|
| Kimi K2.5 | #1 | Moonshot AI | 63.5% |
| MiniMax M3 | #2 | MiniMax | 52.6% |
Average within the visible cohort · descriptive only
| Provider | Models | Average score |
|---|---|---|
| Moonshot AI | 1 | 63.5% |
| MiniMax | 1 | 52.6% |
One best score per model · higher is better
| Rank | Model | Provider | License | Score |
|---|---|---|---|---|
| #1 | Kimi K2.5 kimi-k2.5 | Moonshot AI | — | 63.5% |
| #2 | MiniMax M3 MiniMax-M3 | MiniMax | — | 52.6% |
About PaperBench
An evaluation of agents attempting to replicate machine-learning research from published papers. Results stay tied to the exact model variant and evaluation system. Multiple systems for the same model use the best published score on this page; overall Lumina scoring uses the median of ranking-eligible rows.
Open benchmark source ↗FAQ
An evaluation of agents attempting to replicate machine-learning research from published papers.
Kimi K2.5 by Moonshot AI currently leads with 63.5%.
2 models in the LuminaBench cohort have a qualifying score on this benchmark.
Yes. This family is ranking-weighted at 3.5% of its capability category. Kimi K2.5 is overall #53.
Related
Rolling · 22 results · Weighted 3.5%
SimpleQARolling · 2 results · Weighted 3.5%
Finance Agent v2v2 · 4 results · Not Weighted
BenchLM Knowledge priorbench-align-v5.1 · 33 results · Weighted 1.6%
AA Long Context ReasoningRolling · 96 results · Weighted 1.5%
AA-OmniscienceRolling · 94 results · Weighted 1.5%