Reasoning
Mathematical and scientific problem solving.
Model profile / Z.AI
Z.AI's GLM-5 model tracked via the BenchLM July 2026 leaderboard for capability comparison.
Four areas of evidence
Mathematical and scientific problem solving.
Software engineering and programming.
Factual reliability and instruction following.
Score includes qualified prediction of missing evidence.
Tool use, workflows and professional tasks.
Current index snapshot · 2026-09-16. Each category uses its own scale; scores across categories are not directly comparable. Counts describe published catalogue evidence. A point lead does not establish superiority.
Performance in context
GLM-5 has no current Capabilities rating. Explore the rated models below.
Lumina Capabilities Index
The current Capabilities cohort, with available operational measurements.
Chart loads as you explore
Lumina Capabilities Index
The current Capabilities cohort, with available operational measurements.
Chart loads as you explore
Explore the detail
Score includes qualified prediction of missing evidence.
| Benchmark | Result | Published configuration | Source |
|---|---|---|---|
| IFBench2026 | 72.3% | Exact BenchLM registry variant GLM-5; bulk export does not retain a complete upstream harness configuration.Source-native system | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21source-checked · reference-only |
| IFBench2026 | 72.3% | Exact BenchLM registry variant GLM-5; bulk export does not retain a complete upstream harness configuration.Source-native system | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27source-checked · reference-only |
| IFBench2026 | 72.3% | Exact BenchLM registry variant GLM-5; bulk export does not retain a complete upstream harness configuration.Source-native system | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01source-checked · reference-only |
| AA-Omniscience2026 | 2 index | Exact BenchLM registry variant GLM-5; bulk export does not retain a complete upstream harness configuration.Source-native system | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21source-checked · reference-only |
| AA-Omniscience2026 | 2 index | Exact BenchLM registry variant GLM-5; bulk export does not retain a complete upstream harness configuration.Source-native system | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27source-checked · reference-only |
| AA-Omniscience2026 | 2 index | Exact BenchLM registry variant GLM-5; bulk export does not retain a complete upstream harness configuration.Source-native system | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01source-checked · reference-only |
| SciCode2026 | 46.2% | Exact BenchLM registry variant GLM-5; bulk export does not retain a complete upstream harness configuration.Source-native system | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21source-checked · reference-only |
| SciCode2026 | 46.2% | Exact BenchLM registry variant GLM-5; bulk export does not retain a complete upstream harness configuration.Source-native system | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27source-checked · reference-only |
| SciCode2026 | 46.2% | Exact BenchLM registry variant GLM-5; bulk export does not retain a complete upstream harness configuration.Source-native system | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01source-checked · reference-only |
| AIME25 first-party comparison snapshot2026 | 93.3% | Exact BenchLM registry variant GLM-5; bulk export does not retain a complete upstream harness configuration.Source-native system | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21source-checked · reference-only |
Specialist evidence
Source-native operational evidence for GLM-5. Exact configurations remain separate. Costs are comparable only within the same selected benchmark/evaluation, and these rows never enter Overall Score.
| Evaluation | Exact configuration | Performance | Cost / task | Tokens / task | Execution |
|---|---|---|---|---|---|
| SWE-bench Owner Leaderboards · owner-current · multilingual SWE-bench Owner Leaderboards · checked 2026-08-29 | glm-5-epoch-glm-5 GLM 5 | 69.7% | $0.642 | — | 67.6 calls |