Reasoning
Mathematical and scientific problem solving.
Score and research range
- Full-precision point
- 105.15731858709346
- Range
- 102.54–107.78 · 90% Research Range
- Scope
- Conditional fuller R0 A
Score includes qualified prediction of missing evidence.
Model profile / Google
Google's open-weight Gemma 4 26B A4B instruction model, with 25.2B total and 3.8B active parameters, 256K context, image input, configurable thinking and Apache-2.0 licensing.
Four areas of evidence
Mathematical and scientific problem solving.
Score includes qualified prediction of missing evidence.
Software engineering and programming.
Factual reliability and instruction following.
Tool use, workflows and professional tasks.
Current index snapshot · 2026-09-16. Each category uses its own scale; scores across categories are not directly comparable. Evidence counts describe the retained Capabilities inputs; a family can contain several tasks or configurations. A point lead does not establish superiority.
Performance in context
Gemma 4 26B A4B is highlighted wherever a compatible measurement is available.
Gemma 4 26B A4B: Open-weight Gemma checkpoint has no native Gemini API token-price row; hosted alternatives excluded
Lumina Capabilities Index
The current Capabilities cohort, with this model highlighted.
Chart loads as you explore
Gemma 4 26B A4B: No exact base-profile release and documented 10K output measurement; other effort/protocol/alias records not substituted
Lumina Capabilities Index
The current Capabilities cohort, with this model highlighted.
Chart loads as you explore
Explore the detail
Score uses the admitted direct native evidence under the selected method.
Score includes qualified prediction of missing evidence.
| Benchmark | Result | Published configuration | Source |
|---|---|---|---|
| AA Long Context Reasoningstandard | 61.667% | Gemma 4 26B A4B (Reasoning) (Artificial Analysis completed independent run)Artificial Analysis AA-LCR evaluation | Gemma 4 26B A4B (Reasoning) current model and individual-evaluation page ↗Observed 2026-08-29 · checked 2026-08-29independently-verified · ranking-eligible |
| CritPtstandard | 0% | Gemma 4 26B A4B (Reasoning) (Artificial Analysis completed independent run)Artificial Analysis CritPt evaluation | Gemma 4 26B A4B (Reasoning) current model and individual-evaluation page ↗Observed 2026-08-29 · checked 2026-08-29independently-verified · ranking-eligible |
| GDPval-AA v2v2 | 13.462% | Gemma 4 26B A4B (Reasoning) (Artificial Analysis completed independent run)Artificial Analysis GDPval-AA v2 | Gemma 4 26B A4B (Reasoning) current model and individual-evaluation page ↗Observed 2026-08-29 · checked 2026-08-29independently-verified · ranking-eligible |
| GPQA Diamonddiamond | 79.192% | Gemma 4 26B A4B (Reasoning) (Artificial Analysis completed independent run)Artificial Analysis GPQA Diamond evaluation | Gemma 4 26B A4B (Reasoning) current model and individual-evaluation page ↗Observed 2026-08-29 · checked 2026-08-29independently-verified · ranking-eligible |
| Humanity's Last Examtext-only current | 19.323% | Gemma 4 26B A4B (Reasoning) (Artificial Analysis completed independent run)Artificial Analysis text-only HLE evaluation | Gemma 4 26B A4B (Reasoning) current model and individual-evaluation page ↗Observed 2026-08-29 · checked 2026-08-29independently-verified · ranking-eligible |
| SciCode2024 | 40.046% | Gemma 4 26B A4B (Reasoning) (Artificial Analysis completed independent run)Artificial Analysis SciCode evaluation | Gemma 4 26B A4B (Reasoning) current model and individual-evaluation page ↗Observed 2026-08-29 · checked 2026-08-29independently-verified · reference-only |
| τ³-Banking3 | 11.959% | Gemma 4 26B A4B (Reasoning) (Artificial Analysis completed independent run)Artificial Analysis tau3-Banking evaluation | Gemma 4 26B A4B (Reasoning) current model and individual-evaluation page ↗Observed 2026-08-29 · checked 2026-08-29independently-verified · ranking-eligible |
| Terminal-Bench2.1 | 38.951% | Gemma 4 26B A4B (Reasoning) (Artificial Analysis completed independent run)Artificial Analysis Terminal-Bench v2.1 in e2b | Gemma 4 26B A4B (Reasoning) current model and individual-evaluation page ↗Observed 2026-08-29 · checked 2026-08-29independently-verified · ranking-eligible |
| Artificial Analysis Agentic Index2026 | 10.97 index | Exact BenchLM registry variant Gemma 4 26B A4B; bulk export does not retain a complete upstream harness configuration.Source-native system | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21source-checked · reference-only |
| Artificial Analysis Agentic Index2026 | 10.97 index | Exact BenchLM registry variant Gemma 4 26B A4B; bulk export does not retain a complete upstream harness configuration.Source-native system | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27source-checked · reference-only |
No specialist task observations retained for this model.