Reasoning
Mathematical and scientific problem solving.
Model profile / OpenAI
OpenAI's self-hosted 117B-parameter open-weight model with configurable reasoning, a 131,072-token context and maximum output, released under Apache 2.0.
Four areas of evidence
Mathematical and scientific problem solving.
Software engineering and programming.
Factual reliability and instruction following.
Tool use, workflows and professional tasks.
Score uses the admitted direct native evidence under the selected method.
Current index snapshot · 2026-09-16. Each category uses its own scale; scores across categories are not directly comparable. Evidence counts describe the retained Capabilities inputs; a family can contain several tasks or configurations. A point lead does not establish superiority.
Performance in context
GPT-OSS 120B is highlighted wherever a compatible measurement is available.
GPT-OSS 120B: OpenAI open weights; no first-party hosted API token price
Lumina Capabilities Index
The current Capabilities cohort, with this model highlighted.
Chart loads as you explore
Lumina Capabilities Index
The current Capabilities cohort, with this model highlighted.
Chart loads as you explore
Explore the detail
Score uses the admitted direct native evidence under the selected method.
Score uses the admitted direct native evidence under the selected method.
| Benchmark | Result | Published configuration | Source |
|---|---|---|---|
| Artificial Analysis Agentic Index2026 | 13.17 index | Exact BenchLM registry variant GPT-OSS 120B; bulk export does not retain a complete upstream harness configuration.Source-native system | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21source-checked · reference-only |
| Artificial Analysis Agentic Index2026 | 13.17 index | Exact BenchLM registry variant GPT-OSS 120B; bulk export does not retain a complete upstream harness configuration.Source-native system | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27source-checked · reference-only |
| Artificial Analysis Agentic Index2026 | 13.17 index | Exact BenchLM registry variant GPT-OSS 120B; bulk export does not retain a complete upstream harness configuration.Source-native system | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01source-checked · reference-only |
| Artificial Analysis AIME 20252026 | 93.4% | Exact BenchLM registry variant GPT-OSS 120B; bulk export does not retain a complete upstream harness configuration.Source-native system | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21source-checked · reference-only |
| Artificial Analysis AIME 20252026 | 93.4% | Exact BenchLM registry variant GPT-OSS 120B; bulk export does not retain a complete upstream harness configuration.Source-native system | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27source-checked · reference-only |
| Artificial Analysis AIME 20252026 | 93.4% | Exact BenchLM registry variant GPT-OSS 120B; bulk export does not retain a complete upstream harness configuration.Source-native system | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01source-checked · reference-only |
| Artificial Analysis Coding Index2026 | 30.44 index | Exact BenchLM registry variant GPT-OSS 120B; bulk export does not retain a complete upstream harness configuration.Source-native system | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21source-checked · reference-only |
| Artificial Analysis Coding Index2026 | 30.44 index | Exact BenchLM registry variant GPT-OSS 120B; bulk export does not retain a complete upstream harness configuration.Source-native system | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27source-checked · reference-only |
| Artificial Analysis Coding Index2026 | 30.44 index | Exact BenchLM registry variant GPT-OSS 120B; bulk export does not retain a complete upstream harness configuration.Source-native system | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01source-checked · reference-only |
| Artificial Analysis EnterpriseOps-Gym2026 | 25.5% | Exact BenchLM registry variant GPT-OSS 120B; bulk export does not retain a complete upstream harness configuration.Source-native system | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21source-checked · reference-only |
Specialist evidence
Source-native operational evidence for GPT-OSS 120B. Exact configurations remain separate. Costs are comparable only within the same selected benchmark/evaluation, and these rows never enter Overall Score.
| Evaluation | Exact configuration | Performance | Cost / task | Tokens / task | Execution |
|---|---|---|---|---|---|
| SWE-bench Owner Leaderboards · owner-current · bash-only SWE-bench Owner Leaderboards · checked 2026-08-29 | gpt-oss-120b-high gpt-oss-120b | 26.0% | $0.057 | — | 27.6 calls |
| SWE-bench Owner Leaderboards · owner-current · verified SWE-bench Owner Leaderboards · checked 2026-08-29 | gpt-oss-120b-high mini-SWE-agent + gpt-oss-120b | 26.0% | $0.057 | — | 27.6 calls |