Reasoning
Mathematical and scientific problem solving.
Score and research range
- Full-precision point
- 108.84532769287343
- Range
- 106.23–111.46 · 90% Research Range
- Scope
- Conditional fuller R0 A
Score includes qualified prediction of missing evidence.
Model profile / Meituan
Meituan's open-weight 1.6T-A48B Mixture-of-Experts model with LongCat Sparse Attention, a 1,000,000-token context window, MIT licence, and official in-house agent and coding scores.
Four areas of evidence
Mathematical and scientific problem solving.
Score includes qualified prediction of missing evidence.
Software engineering and programming.
Factual reliability and instruction following.
Tool use, workflows and professional tasks.
Current index snapshot · 2026-09-16. Each category uses its own scale; scores across categories are not directly comparable. Evidence counts describe the retained Capabilities inputs; a family can contain several tasks or configurations. A point lead does not establish superiority.
Performance in context
LongCat-2.0 is highlighted wherever a compatible measurement is available.
Lumina Capabilities Index
The current Capabilities cohort, with this model highlighted.
Chart loads as you explore
LongCat-2.0: No exact base-profile release and documented 10K output measurement; other effort/protocol/alias records not substituted
Lumina Capabilities Index
The current Capabilities cohort, with this model highlighted.
Chart loads as you explore
Explore the detail
Score uses the admitted direct native evidence under the selected method.
Score includes qualified prediction of missing evidence.
| Benchmark | Result | Published configuration | Source |
|---|---|---|---|
| AA Long Context Reasoningstandard | 62.667% | LongCat 2.0 (Artificial Analysis independent run)Artificial Analysis AA-LCR evaluation | LongCat 2.0 individual evaluations ↗Observed 2026-08-29 · checked 2026-08-29independently-verified · ranking-eligible |
| CritPtstandard | 2.571% | LongCat 2.0 (Artificial Analysis independent run)Artificial Analysis CritPt evaluation | LongCat 2.0 individual evaluations ↗Observed 2026-08-29 · checked 2026-08-29independently-verified · ranking-eligible |
| GDPval-AA v2v2 | 26.61% | LongCat 2.0 (Artificial Analysis independent run)Artificial Analysis GDPval-AA v2 | LongCat 2.0 individual evaluations ↗Observed 2026-08-29 · checked 2026-08-29independently-verified · ranking-eligible |
| GPQA Diamonddiamond | 77.98% | LongCat 2.0 (Artificial Analysis independent run)Artificial Analysis GPQA Diamond evaluation | LongCat 2.0 individual evaluations ↗Observed 2026-08-29 · checked 2026-08-29independently-verified · ranking-eligible |
| Humanity's Last Examtext-only current | 33.689% | LongCat 2.0 (Artificial Analysis independent run)Artificial Analysis text-only HLE evaluation | LongCat 2.0 individual evaluations ↗Observed 2026-08-29 · checked 2026-08-29independently-verified · ranking-eligible |
| SciCode2024 | 35.417% | LongCat 2.0 (Artificial Analysis independent run)Artificial Analysis SciCode evaluation | LongCat 2.0 individual evaluations ↗Observed 2026-08-29 · checked 2026-08-29independently-verified · reference-only |
| τ³-Banking3 | 13.196% | LongCat 2.0 (Artificial Analysis independent run)Artificial Analysis tau3-Banking evaluation | LongCat 2.0 individual evaluations ↗Observed 2026-08-29 · checked 2026-08-29independently-verified · ranking-eligible |
| Terminal-Bench2.1 | 50.187% | LongCat 2.0 (Artificial Analysis independent run)Artificial Analysis Terminal-Bench v2.1 in e2b | LongCat 2.0 individual evaluations ↗Observed 2026-08-29 · checked 2026-08-29independently-verified · ranking-eligible |
| IMOAnswerBench2026 | 81.8% | LongCat-2.0 (in-house harness)LongCat unified in-house harness | LongCat-2.0 model card ↗Observed 2026-06-30 · checked 2026-08-15provider-reported · reference-only |
| SWE-bench Multilingual2025 | 77.3% | LongCat-2.0 (in-house harness)LongCat unified in-house harness | LongCat-2.0 model card ↗Observed 2026-06-30 · checked 2026-08-15provider-reported · reference-only |