Reasoning
Mathematical and scientific problem solving.
Score and research range
- Full-precision point
- 114.85460394115246
- Range
- 112.24–117.47 · 90% Research Range
- Scope
- Conditional fuller R0 A
Score includes qualified prediction of missing evidence.
Model profile / Meta
Meta Muse Spark 1.1 frontier API model with configurable reasoning, 1M context, and strong agentic scores on provider and independent evaluations.
Four areas of evidence
Mathematical and scientific problem solving.
Score includes qualified prediction of missing evidence.
Software engineering and programming.
Score uses the admitted direct native evidence under the selected method.
Factual reliability and instruction following.
Score uses the admitted direct native evidence under the selected method.
Tool use, workflows and professional tasks.
Score uses the admitted direct native evidence under the selected method.
Current index snapshot · 2026-09-16. Each category uses its own scale; scores across categories are not directly comparable. Evidence counts describe the retained Capabilities inputs; a family can contain several tasks or configurations. A point lead does not establish superiority.
Performance in context
Muse Spark 1.1 is highlighted wherever a compatible measurement is available.
Muse Spark 1.1: First-party page did not expose a verifiable checkpoint-specific token price; no AA/provider-median substitute
Lumina Capabilities Index
The current Capabilities cohort, with this model highlighted.
Chart loads as you explore
Muse Spark 1.1: No exact base-profile release and documented 10K output measurement; other effort/protocol/alias records not substituted
Lumina Capabilities Index
The current Capabilities cohort, with this model highlighted.
Chart loads as you explore
Explore the detail
Score uses the admitted direct native evidence under the selected method.
Score uses the admitted direct native evidence under the selected method.
Score includes qualified prediction of missing evidence.
Score uses the admitted direct native evidence under the selected method.
Score uses the admitted direct native evidence under the selected method.
| Benchmark | Result | Published configuration | Source |
|---|---|---|---|
| AA Long Context Reasoningstandard | 81.333% | Muse Spark 1.1 (xhigh; Artificial Analysis independent run)Artificial Analysis AA-LCR evaluation | Muse Spark 1.1 individual evaluations ↗Observed 2026-08-27 · checked 2026-08-27independently-verified · ranking-eligible |
| CritPtstandard | 15.143% | Muse Spark 1.1 (xhigh; Artificial Analysis independent run)Artificial Analysis CritPt evaluation | Muse Spark 1.1 individual evaluations ↗Observed 2026-08-27 · checked 2026-08-27independently-verified · ranking-eligible |
| GDPval-AA v2v2 | 43.66% | Muse Spark 1.1 (xhigh; Artificial Analysis independent run)Artificial Analysis GDPval-AA v2 using Stirrup | Muse Spark 1.1 individual evaluations ↗Observed 2026-08-27 · checked 2026-08-27independently-verified · ranking-eligible |
| GPQA Diamonddiamond | 89.798% | Muse Spark 1.1 (xhigh; Artificial Analysis independent run)Artificial Analysis GPQA Diamond evaluation | Muse Spark 1.1 individual evaluations ↗Observed 2026-08-27 · checked 2026-08-27independently-verified · ranking-eligible |
| Humanity's Last Examtext-only current | 46.2% | Muse Spark 1.1 (xhigh; Artificial Analysis independent run)Artificial Analysis text-only Humanity's Last Exam evaluation | Muse Spark 1.1 individual evaluations ↗Observed 2026-08-27 · checked 2026-08-27independently-verified · ranking-eligible |
| SciCoderolling | 58.218% | Muse Spark 1.1 (xhigh; Artificial Analysis independent run)Artificial Analysis SciCode evaluation | Muse Spark 1.1 individual evaluations ↗Observed 2026-08-27 · checked 2026-08-27independently-verified · reference-only |
| τ³-Banking3 | 31.753% | Muse Spark 1.1 (xhigh; Artificial Analysis independent run)Artificial Analysis tau3-Banking evaluation | Muse Spark 1.1 individual evaluations ↗Observed 2026-08-27 · checked 2026-08-27independently-verified · ranking-eligible |
| Terminal-Bench2.1 | 77.903% | Muse Spark 1.1 (xhigh; Artificial Analysis independent run)Artificial Analysis Terminal-Bench v2.1 in e2b | Muse Spark 1.1 individual evaluations ↗Observed 2026-08-27 · checked 2026-08-27independently-verified · ranking-eligible |
| Artificial Analysis Agentic Index2026 | 37.54 index | Exact BenchLM registry variant Muse Spark 1.1; bulk export does not retain a complete upstream harness configuration.Source-native system | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21source-checked · reference-only |
| Artificial Analysis Agentic Index2026 | 37.54 index | Exact BenchLM registry variant Muse Spark 1.1; bulk export does not retain a complete upstream harness configuration.Source-native system | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27source-checked · reference-only |
Specialist evidence
Source-native operational evidence for Muse Spark 1.1. Exact configurations remain separate. Costs are comparable only within the same selected benchmark/evaluation, and these rows never enter Overall Score.
| Evaluation | Exact configuration | Performance | Cost / task | Tokens / task | Execution |
|---|---|---|---|---|---|
| Artificial Analysis Coding Agents · 1.4 · composite Artificial Analysis Coding Agents · checked 2026-08-29 | muse-spark-1-1-xhigh Opencode - Muse Spark 1.1 (xhigh) | 54.9 | $1.44 | 12289527 | 55.5 steps |
| Terminal-Bench 2.1 · 2.1 · verified Terminal-Bench 2.1 · checked 2026-08-29 | muse-spark-1-1-xhigh mini-SWE-agent; reasoning xhigh; Terminal-Bench 2.1 verified submission. | 76.2% | $2.23 | — | — |