Reasoning
Mathematical and scientific problem solving.
Score and research range
- Full-precision point
- 108.6333752884098
- Range
- 106.01–111.25 · 90% Research Range
- Scope
- Conditional fuller R0 A
Score includes qualified prediction of missing evidence.
Model profile / OpenAI
OpenAI's GPT-5 coding variant for agentic Codex workflows through the Responses API, with a 400K context window and 128K maximum output.
Four areas of evidence
Mathematical and scientific problem solving.
Score includes qualified prediction of missing evidence.
Software engineering and programming.
Factual reliability and instruction following.
Tool use, workflows and professional tasks.
Current index snapshot · 2026-09-16. Each category uses its own scale; scores across categories are not directly comparable. Counts describe published catalogue evidence. A point lead does not establish superiority.
Performance in context
GPT-5 Codex has no current Capabilities rating. Explore the rated models below.
Lumina Capabilities Index
The current Capabilities cohort, with available operational measurements.
Chart loads as you explore
Lumina Capabilities Index
The current Capabilities cohort, with available operational measurements.
Chart loads as you explore
Explore the detail
Score includes qualified prediction of missing evidence.
| Benchmark | Result | Published configuration | Source |
|---|---|---|---|
| AA Long Context ReasoningSource-native version | 69% | GPT-5 Codex (high)Artificial Analysis independent evaluation | AA Long Context Reasoning Leaderboard ↗Observed 2026-07-15 · checked 2026-07-15source-checked · reference-only |
| AA Long Context ReasoningSource-native version | 69% | GPT-5 Codex (high)Artificial Analysis independent evaluation | AA Long Context Reasoning Leaderboard ↗Observed 2026-07-15 · checked 2026-07-15source-checked · reference-only |
| AA-OmniscienceSource-native version | -6.767 index | GPT-5 Codex (high)Artificial Analysis independent evaluation | AA-Omniscience Leaderboard ↗Observed 2026-07-15 · checked 2026-07-15source-checked · reference-only |
| AA-OmniscienceSource-native version | -6.767 index | GPT-5 Codex (high)Artificial Analysis independent evaluation | AA-Omniscience Leaderboard ↗Observed 2026-07-15 · checked 2026-07-15source-checked · reference-only |
| CritPtSource-native version | 5.143% | GPT-5 Codex (high)Artificial Analysis independent evaluation | CritPt Leaderboard ↗Observed 2026-07-15 · checked 2026-07-15source-checked · ranking-eligible |
| CritPtSource-native version | 5.143% | GPT-5 Codex (high)Artificial Analysis independent evaluation | CritPt Leaderboard ↗Observed 2026-07-15 · checked 2026-07-15source-checked · reference-only |
| τ²-Bench TelecomSource-native version | 86.842% | GPT-5 Codex (high)Artificial Analysis independent evaluation | τ²-Bench Telecom Leaderboard ↗Observed 2026-07-15 · checked 2026-07-15source-checked · reference-only |
| τ²-Bench TelecomSource-native version | 86.842% | GPT-5 Codex (high)Artificial Analysis independent evaluation | τ²-Bench Telecom Leaderboard ↗Observed 2026-07-15 · checked 2026-07-15source-checked · reference-only |
| Terminal-Bench HardSource-native version | 37.879% | GPT-5 Codex (high)Artificial Analysis independent evaluation | Terminal-Bench Hard Leaderboard ↗Observed 2026-07-15 · checked 2026-07-15source-checked · reference-only |
| Terminal-Bench HardSource-native version | 37.879% | GPT-5 Codex (high)Artificial Analysis independent evaluation | Terminal-Bench Hard Leaderboard ↗Observed 2026-07-15 · checked 2026-07-15source-checked · reference-only |
No specialist task observations retained for this model.