| AA Long Context Reasoningstandard | 74.333% | Claude Opus 4.6 (Adaptive Reasoning, Max Effort) (Artificial Analysis independent run)Artificial Analysis AA-LCR evaluation | Claude Opus 4.6 (Adaptive Reasoning, Max Effort) individual evaluations ↗Observed 2026-08-29 · checked 2026-08-29independently-verified · reference-only |
|---|
| CritPtstandard | 12.571% | Claude Opus 4.6 (Adaptive Reasoning, Max Effort) (Artificial Analysis independent run)Artificial Analysis CritPt evaluation | Claude Opus 4.6 (Adaptive Reasoning, Max Effort) individual evaluations ↗Observed 2026-08-29 · checked 2026-08-29independently-verified · reference-only |
|---|
| GPQA Diamonddiamond | 89.596% | Claude Opus 4.6 (Adaptive Reasoning, Max Effort) (Artificial Analysis independent run)Artificial Analysis GPQA Diamond evaluation | Claude Opus 4.6 (Adaptive Reasoning, Max Effort) individual evaluations ↗Observed 2026-08-29 · checked 2026-08-29independently-verified · reference-only |
|---|
| Humanity's Last Examtext-only current | 39.944% | Claude Opus 4.6 (Adaptive Reasoning, Max Effort) (Artificial Analysis independent run)Artificial Analysis text-only HLE evaluation | Claude Opus 4.6 (Adaptive Reasoning, Max Effort) individual evaluations ↗Observed 2026-08-29 · checked 2026-08-29independently-verified · reference-only |
|---|
| SciCode2024 | 51.852% | Claude Opus 4.6 (Adaptive Reasoning, Max Effort) (Artificial Analysis independent run)Artificial Analysis SciCode evaluation | Claude Opus 4.6 (Adaptive Reasoning, Max Effort) individual evaluations ↗Observed 2026-08-29 · checked 2026-08-29independently-verified · reference-only |
|---|
| IFBench2026 | 44.6% | Exact BenchLM registry variant Claude Opus 4.6; bulk export does not retain a complete upstream harness configuration.Source-native system | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21source-checked · reference-only |
|---|
| IFBench2026 | 44.6% | Exact BenchLM registry variant Claude Opus 4.6; bulk export does not retain a complete upstream harness configuration.Source-native system | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27source-checked · reference-only |
|---|
| IFBench2026 | 44.6% | Exact BenchLM registry variant Claude Opus 4.6; bulk export does not retain a complete upstream harness configuration.Source-native system | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01source-checked · reference-only |
|---|
| MMMU-Pro2026 | 72.5% | Exact BenchLM registry variant Claude Opus 4.6; bulk export does not retain a complete upstream harness configuration.Source-native system | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21source-checked · reference-only |
|---|
| MMMU-Pro2026 | 72.5% | Exact BenchLM registry variant Claude Opus 4.6; bulk export does not retain a complete upstream harness configuration.Source-native system | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27source-checked · reference-only |
|---|