A task with its own rules.
Harder agentic terminal evaluation suite spanning software engineering, systems administration, and data processing tasks.
- Organisation
- Laude Institute / Artificial Analysis
- Version
- Current / rolling
Agents / Benchmark profile
Harder agentic terminal evaluation suite spanning software engineering, systems administration, and data processing tasks.
Observed results
20 ranked models · higher is better · labels show rank and score
| Model | Rank | Model ID | Provider | Score | Unit |
|---|---|---|---|---|---|
| Kimi K3 | #1 | kimi-k3 | Moonshot AI | 85 | percent |
| GPT-5.6 Sol | #2 | gpt-5-6-sol | OpenAI | 65.9091 | percent |
| Claude Fable 5 | #3 | claude-fable-5 | Anthropic | 62.9 | percent |
| GPT-5.5 | #4 | gpt-5-5 | OpenAI | 60.6061 | percent |
| Claude Opus 4.8 | #5 | claude-opus-4-8 | Anthropic | 58.3333 | percent |
| GPT-5.6 Terra | #6 | gpt-5-6-terra | OpenAI | 57.6 | percent |
| GPT-5.4 | #7 | gpt-5-4 | OpenAI | 57.5758 | percent |
| Gemini 3.1 Pro Preview | #8 | gemini-3-1-pro-preview | 53.7879 | percent | |
| Claude Sonnet 4.6 | #9 | claude-sonnet-4-6 | Anthropic | 53.0303 | percent |
| GPT-5.3-Codex | #10 | gpt-5-3-codex | OpenAI | 53.0303 | percent |
| GPT-5.4 mini | #11 | gpt-5-4-mini | OpenAI | 52.2727 | percent |
| Claude Opus 4.7 | #12 | claude-opus-4-7 | Anthropic | 51.5152 | percent |
| GLM-5.2 | #13 | glm-5-2 | Z.AI | 50.8 | percent |
| Qwen3.7-Max | #14 | qwen-3-7-max | Alibaba Cloud | 50.8 | percent |
| Claude Opus 4.5 | #15 | claude-opus-4-5 | Anthropic | 46.9697 | percent |
| GPT-5.2 | #16 | gpt-5-2 | OpenAI | 46.9697 | percent |
| Qwen3.7-Plus | #17 | qwen-3-7-plus | Alibaba Cloud | 46.9697 | percent |
| Claude Opus 4.6 | #18 | claude-opus-4-6 | Anthropic | 46.2121 | percent |
| DeepSeek V4 Pro | #19 | deepseek-v4-pro | DeepSeek | 46.2121 | percent |
| GPT-5.1 | #20 | gpt-5-1 | OpenAI | 45.4545 | percent |
| Muse Spark | #21 | muse-spark | Meta | 45.4545 | percent |
| Kimi K2.7 Code | #22 | kimi-k2-7-code | Moonshot AI | 44.697 | percent |
| Kimi K2.6 | #23 | kimi-k2-6 | Moonshot AI | 43.9394 | percent |
| Qwen3.6 Max | #24 | qwen3-6-max | Alibaba Cloud | 43.9394 | percent |
| Qwen3.6 Plus | #25 | qwen3-6-plus | Alibaba Cloud | 43.9394 | percent |
| MiMo-V2.5-Pro | #26 | mimo-v2-5-pro | Xiaomi | 43.2 | percent |
| GLM-5 | #27 | glm-5 | Z.AI | 43.1818 | percent |
| GLM-5.1 | #28 | glm-5-1 | Z.AI | 43.1818 | percent |
| GPT-5.4 nano | #29 | gpt-5-4-nano | OpenAI | 42.4242 | percent |
| MiniMax M3 | #30 | minimax-m3 | MiniMax | 42.4242 | percent |
| MiMo-V2.5 | #31 | mimo-v2-5 | Xiaomi | 41.6667 | percent |
| Gemini 3.5 Flash | #32 | gemini-3-5-flash | 40.9091 | percent | |
| MiMo-V2-Pro | #33 | mimo-v2-pro | Xiaomi | 40.9091 | percent |
| Qwen3.5 397B A17B | #34 | qwen3-5-397b | Alibaba Cloud | 40.9091 | percent |
| MiniMax M2.7 | #35 | minimax-m2-7 | MiniMax | 39.3939 | percent |
| Gemini 3 Flash | #36 | gemini-3-flash | 38.6364 | percent | |
| GPT-5 Codex | #37 | gpt-5-codex | OpenAI | 37.8788 | percent |
| Grok 4 | #38 | grok-4 | xAI | 37.8788 | percent |
| Grok 4.20 | #39 | grok-4-20 | xAI | 37.8788 | percent |
| Grok 4.3 | #40 | grok-4-3 | xAI | 37.8788 | percent |
| GPT-5.2 Codex | #41 | gpt-5-2-codex | OpenAI | 37.1212 | percent |
| o3 | #42 | o3 | OpenAI | 37.1212 | percent |
| Gemma 4 31B | #43 | gemma-4-31b | 36.4 | percent | |
| Nemotron 3 Ultra | #44 | nemotron-3-ultra | NVIDIA | 36.4 | percent |
| Claude Sonnet 4.5 | #45 | claude-sonnet-4-5 | Anthropic | 35.6061 | percent |
| DeepSeek V3.2 | #46 | deepseek-v3-2 | DeepSeek | 35.6061 | percent |
| DeepSeek V4 Flash | #47 | deepseek-v4-flash | DeepSeek | 35.6061 | percent |
| Kimi K2.5 | #48 | kimi-k2-5 | Moonshot AI | 34.8485 | percent |
| MiMo-V2-Omni | #49 | mimo-v2-omni | Xiaomi | 34.8485 | percent |
| MiniMax M2.5 | #50 | minimax-m2-5 | MiniMax | 34.8485 | percent |
| Qwen3.6 27B | #51 | qwen3-6-27b | Alibaba Cloud | 34.8485 | percent |
| Claude Opus 4.1 | #52 | claude-opus-4-1 | Anthropic | 34.3352 | percent |
| GLM-5-Turbo | #53 | glm-5-turbo | Z.AI | 33.3333 | percent |
| Mistral Medium 3.5 | #54 | mistral-medium-3-5 | Mistral AI | 33.3333 | percent |
| Mistral Medium 3.5 128B | #55 | mistral-medium-3-5-128b | Mistral AI | 33.3 | percent |
| GLM-5V-Turbo | #56 | glm-5v-turbo | Z.AI | 32.5758 | percent |
| GPT-5 | #57 | gpt-5 | OpenAI | 32.5758 | percent |
| Qwen3.5 27B | #58 | qwen3-5-27b | Alibaba Cloud | 32.5758 | percent |
| GLM-4.7 | #59 | glm-4-7 | Z.AI | 31.8182 | percent |
| Claude Opus 4 | #60 | claude-opus-4 | Anthropic | 31.0606 | percent |
| Claude Sonnet 4 | #61 | claude-sonnet-4 | Anthropic | 31.0606 | percent |
| Kimi K2 Thinking | #62 | kimi-k2-thinking | Moonshot AI | 31.0606 | percent |
| Ling 2.6 1T | #63 | ling-2-6-1t | InclusionAI | 31.0606 | percent |
| Qwen3.5 122B | #64 | qwen3-5-122b | Alibaba Cloud | 31.0606 | percent |
| DeepSeek V3.1 Terminus | #65 | deepseek-v3-1-terminus | DeepSeek | 30.303 | percent |
| GPT-5 mini | #66 | gpt-5-mini | OpenAI | 28.7879 | percent |
| MiniMax M2.1 | #67 | minimax-m2-1 | MiniMax | 28.7879 | percent |
| Nemotron 3 Super | #68 | nemotron-3-super | NVIDIA | 28.7879 | percent |
| Solar Open 2 250B | #69 | solar-open2-250b | Upstage | 28.3 | percent |
| Claude Haiku 4.5 | #70 | claude-haiku-4-5 | Anthropic | 27.2727 | percent |
| Gemini 2.5 Pro | #71 | gemini-2-5-pro | 26.5152 | percent | |
| Qwen3.5 35B | #72 | qwen3-5-35b | Alibaba Cloud | 26.5152 | percent |
| MiniMax M2 | #73 | minimax-m2 | MiniMax | 25.7576 | percent |
| Command A+ | #74 | command-a-plus | Cohere | 25 | percent |
| GLM-4.6 | #75 | glm-4-6 | Z.AI | 25 | percent |
| Gemini 3.1 Flash-Lite | #76 | gemini-3-1-flash-lite | 24.2424 | percent | |
| Nova 2 Pro | #77 | nova-2-pro | Amazon | 24.2424 | percent |
| Qwen3 Max | #78 | qwen3-max | Alibaba Cloud | 24.2424 | percent |
| GPT-OSS 120B | #79 | gpt-oss-120b | OpenAI | 23.5 | percent |
| Kimi K2 0905 | #80 | kimi-k2-0905 | Moonshot AI | 23.4848 | percent |
| Claude Sonnet 3.7 | #81 | claude-sonnet-3-7 | Anthropic | 21.2121 | percent |
| Qwen3.5 Omni Plus | #82 | qwen3-5-omni-plus | Alibaba Cloud | 21.2121 | percent |
| Grok 4 Fast | #83 | grok-4-fast | xAI | 18.9394 | percent |
| Gemma 4 12B Unified | #84 | gemma-4-12b | 18.1818 | percent | |
| Grok 3 mini | #85 | grok-3-mini | xAI | 17.4242 | percent |
| Grok Code Fast 1 | #86 | grok-code-fast-1 | xAI | 17.4242 | percent |
| Mistral Small 4 | #87 | mistral-small-4 | Mistral AI | 17.4242 | percent |
| Gemini 2.5 Flash | #88 | gemini-2-5-flash | 16.6667 | percent | |
| Mistral Large 3 | #89 | mistral-large-3 | Mistral AI | 15.9091 | percent |
| o4-mini | #90 | o4-mini | OpenAI | 15.1515 | percent |
| Gemma 4 26B | #91 | gemma-4-26b | 13.64 | percent | |
| o1 | #92 | o1 | OpenAI | 12.8788 | percent |
One best compatible score per canonical product · higher is better
| Rank | Model | Weights | Evidence | Score |
|---|---|---|---|---|
| 01 | Kimi K3Moonshot AI | Open weights | Public reference | 85% |
| 02 | GPT-5.6 SolOpenAI | Closed weights | Public reference | 65.9% |
| 03 | Claude Fable 5Anthropic | Closed weights | Source carrier | 62.9% |
| 04 | GPT-5.5OpenAI | Closed weights | Public reference | 60.6% |
| 05 | Claude Opus 4.8Anthropic | Closed weights | Public reference | 58.3% |
| 06 | GPT-5.6 TerraOpenAI | Closed weights | Source carrier | 57.6% |
| 07 | GPT-5.4OpenAI | Closed weights | Public reference | 57.6% |
| 08 | Gemini 3.1 Pro PreviewGoogle | Closed weights | Public reference | 53.8% |
| 09 | Claude Sonnet 4.6Anthropic | Closed weights | Public reference | 53.0% |
| 10 | GPT-5.3-CodexOpenAI | Closed weights | Public reference | 53.0% |
Each distinct published measurement is retained. Same-snapshot canonical and configuration projections appear once.
| Published model / configuration | Score | Evidence & protocol | Source & dates |
|---|---|---|---|
MiMo-V2.5 as published by UpstageMiMo-V2.5Exact identitymimo-v2-5-solar-open2-unspecified Canonical product: mimo-v2-5 | 41.7% | Relative comparisonprovider-reportedVersion & systemHard Solar Open 2 card English table | Solar Open 2 250B model cardObserved Checked |
DeepSeek V4 Flash max as published by UpstageDeepSeek V4 FlashExact identitydeepseek-v4-flash-max Canonical product: deepseek-v4-flash | 34.1% | Relative comparisonprovider-reportedVersion & systemHard Solar Open 2 card English table | Solar Open 2 250B model cardObserved Checked |
Mistral Medium 3.5 as published by UpstageMistral Medium 3.5Exact identitymistral-medium-3-5-high Canonical product: mistral-medium-3-5 | 33.3% | Relative comparisonprovider-reportedVersion & systemHard Solar Open 2 card English table | Solar Open 2 250B model cardObserved Checked |
Solar Open 2 250B (high)Solar Open 2 250BExact identitysolar-open2-250b-high Canonical product: solar-open2-250b | 28.3% | Public referenceprovider-reportedVersion & systemHard Solar Open 2 card English table | Solar Open 2 250B model cardObserved Checked |
Command A+ as published by UpstageCommand A+Exact identitycommand-a-plus-solar-open2-unspecified Canonical product: command-a-plus | 25% | Relative comparisonprovider-reportedVersion & systemHard Solar Open 2 card English table | Solar Open 2 250B model cardObserved Checked |
Solar Open 100B as published by UpstageSolar Open 100B (Reasoning)Exact identitysolar-open-100b-reasoning-high Canonical product: solar-open-100b-reasoning | 2.3% | Relative comparisonprovider-reportedVersion & systemHard Solar Open 2 card English table | Solar Open 2 250B model cardObserved Checked |
Exact BenchLM registry variant GPT-5.6 Sol; bulk export does not retain a complete upstream harness configuration.GPT-5.6 SolExact identitySource label without a registered configuration ID Canonical product: gpt-5-6-sol | 65.9% | Source carriersource-checkedVersion & system2026 Source-native system | BenchLM public datasets — 2026-08-01Observed Checked |
Exact BenchLM registry variant Claude Fable 5; bulk export does not retain a complete upstream harness configuration.Claude Fable 5Exact identitySource label without a registered configuration ID Canonical product: claude-fable-5 | 62.9% | Source carriersource-checkedVersion & system2026 Source-native system | BenchLM public datasets — 2026-08-01Observed Checked |
Exact BenchLM registry variant Claude Opus 4.8; bulk export does not retain a complete upstream harness configuration.Claude Opus 4.8Exact identitySource label without a registered configuration ID Canonical product: claude-opus-4-8 | 58.3% | Source carriersource-checkedVersion & system2026 Source-native system | BenchLM public datasets — 2026-08-01Observed Checked |
Exact BenchLM registry variant GPT-5.6 Terra; bulk export does not retain a complete upstream harness configuration.GPT-5.6 TerraExact identitySource label without a registered configuration ID Canonical product: gpt-5-6-terra | 57.6% | Source carriersource-checkedVersion & system2026 Source-native system | BenchLM public datasets — 2026-08-01Observed Checked |
From result to context
Harder agentic terminal evaluation suite spanning software engineering, systems administration, and data processing tasks.
The summary shows one best compatible source result per canonical product. Versions, effort settings and execution systems remain attached to the underlying records.
Scores from different versions or harnesses may not be interchangeable. The published source rows preserve those distinctions and their original units. A source result is not automatically an input to a current index.
This catalogue entry preserves available source evidence. Current indices admit only their specifically reviewed tracks and configurations.
Contamination risk: Low. Lifecycle: Active.