A task with its own rules.
Instruction-following evaluation measuring constraint and format reliability.
- Organisation
- IFBench
- Version
- Current / rolling
Instruction Following / Benchmark profile
Instruction-following evaluation measuring constraint and format reliability.
Observed results
20 ranked models · higher is better · labels show rank and score
| Model | Rank | Model ID | Provider | Score | Unit |
|---|---|---|---|---|---|
| MiniMax M3 | #1 | minimax-m3 | MiniMax | 82.9 | percent |
| Qwen3.8 Max | #2 | qwen-3-8-max | Alibaba Cloud | 82.8 | percent |
| Nemotron 3 Ultra | #3 | nemotron-3-ultra | NVIDIA | 81.4 | percent |
| Grok 4.3 | #4 | grok-4-3 | xAI | 81.3 | percent |
| Qwen3.8-Flash-Next | #5 | qwen-3-8-flash-next | Alibaba Cloud | 81.3 | percent |
| Grok 4.20 | #6 | grok-4-20 | xAI | 81.2245 | percent |
| Qwen3.7-Max | #7 | qwen-3-7-max | Alibaba Cloud | 80.5 | percent |
| Solar Open 2 250B | #8 | solar-open2-250b | Upstage | 80 | percent |
| MiMo-V2.5-Pro | #9 | mimo-v2-5-pro | Xiaomi | 79.9 | percent |
| Qwen3.8-27B | #10 | qwen-3-8-27b | Alibaba Cloud | 79.5 | percent |
| DeepSeek V4 Flash | #11 | deepseek-v4-flash | DeepSeek | 79.2 | percent |
| Qwen3.7-Plus | #12 | qwen-3-7-plus | Alibaba Cloud | 79.1 | percent |
| Nova 2 Pro | #13 | nova-2-pro | Amazon | 79.0476 | percent |
| Qwen3.5 397B A17B | #14 | qwen3-5-397b | Alibaba Cloud | 78.8 | percent |
| GPT-5.2 Codex | #15 | gpt-5-2-codex | OpenAI | 77.619 | percent |
| Gemini 3.1 Flash-Lite | #16 | gemini-3-1-flash-lite | 77.2 | percent | |
| Gemini 3.1 Pro Preview | #17 | gemini-3-1-pro-preview | 77.1 | percent | |
| Muse Glimmer 30B | #18 | muse-glimmer-30b | Meta | 77 | percent |
| Qwen 3.6 Max (preview) | #19 | qwen3-6-max-preview | Alibaba Cloud | 76.6 | percent |
| Qwen3.6 Max | #20 | qwen3-6-max | Alibaba Cloud | 76.5986 | percent |
| DeepSeek V4 Pro | #21 | deepseek-v4-pro | DeepSeek | 76.5 | percent |
| Gemini 3.5 Flash | #22 | gemini-3-5-flash | 76.3 | percent | |
| GLM-5.1 | #23 | glm-5-1 | Z.AI | 76.3 | percent |
| Kimi K2.6 | #24 | kimi-k2-6 | Moonshot AI | 76 | percent |
| GPT-5.4 nano | #25 | gpt-5-4-nano | OpenAI | 75.9 | percent |
| GPT-5.5 | #26 | gpt-5-5 | OpenAI | 75.9 | percent |
| Muse Spark | #27 | muse-spark | Meta | 75.9 | percent |
| Qwen3.6 Plus | #28 | qwen3-6-plus | Alibaba Cloud | 75.8 | percent |
| Qwen3.5 122B | #29 | qwen3-5-122b | Alibaba Cloud | 75.7143 | percent |
| MiniMax M2.7 | #30 | minimax-m2-7 | MiniMax | 75.7 | percent |
| Qwen3.5-122B-A10B | #31 | qwen3-5-122b-a10b | Alibaba Cloud | 75.7 | percent |
| Gemma 4 31B | #32 | gemma-4-31b | 75.6 | percent | |
| Qwen3.5 27B | #33 | qwen3-5-27b | Alibaba Cloud | 75.6 | percent |
| GPT-5.2 | #34 | gpt-5-2 | OpenAI | 75.4422 | percent |
| GPT-5.3-Codex | #35 | gpt-5-3-codex | OpenAI | 75.4 | percent |
| GPT-5 Codex | #36 | gpt-5-codex | OpenAI | 74.1497 | percent |
| Command A+ | #37 | command-a-plus | Cohere | 73.9456 | percent |
| GPT-5.4 | #38 | gpt-5-4 | OpenAI | 73.9 | percent |
| Gemma 4 12B Unified | #39 | gemma-4-12b | 73.5374 | percent | |
| GLM-5.2 | #40 | glm-5-2 | Z.AI | 73.3 | percent |
| GPT-5.4 mini | #41 | gpt-5-4-mini | OpenAI | 73.3 | percent |
| GLM-5-Turbo | #42 | glm-5-turbo | Z.AI | 73.2 | percent |
| GPT-5 | #43 | gpt-5 | OpenAI | 73.1 | percent |
| GPT-5.1 | #44 | gpt-5-1 | OpenAI | 72.9 | percent |
| GPT-5.6 Sol | #45 | gpt-5-6-sol | OpenAI | 72.7 | percent |
| Qwen3.5 35B | #46 | qwen3-5-35b | Alibaba Cloud | 72.517 | percent |
| Qwen3.5-35B-A3B | #47 | qwen3-5-35b-a3b | Alibaba Cloud | 72.5 | percent |
| Gemma 4 26B | #48 | gemma-4-26b | 72.449 | percent | |
| Gemma 4 26B A4B | #49 | gemma-4-26b-a4b | 72.4 | percent | |
| MiniMax M2 | #50 | minimax-m2 | MiniMax | 72.3129 | percent |
| GLM-5 | #51 | glm-5 | Z.AI | 72.3 | percent |
| NVIDIA Nemotron 3.5 Lightning 30B-A3B | #52 | nemotron-3-5-lightning-30b-a3b | NVIDIA | 71.88 | percent |
| MiniMax M2.5 | #53 | minimax-m2-5 | MiniMax | 71.6327 | percent |
| Nemotron 3 Super | #54 | nemotron-3-super | NVIDIA | 71.4966 | percent |
| o3 | #55 | o3 | OpenAI | 71.4286 | percent |
| GPT-5.6 Terra | #56 | gpt-5-6-terra | OpenAI | 71.2 | percent |
| GPT-5 mini | #57 | gpt-5-mini | OpenAI | 71.1565 | percent |
| Nemotron 3 Nano 30B | #58 | nemotron-3-nano-30b | NVIDIA | 71.1 | percent |
| Qwen3 Max | #59 | qwen3-max | Alibaba Cloud | 70.7483 | percent |
| o1 | #60 | o1 | OpenAI | 70.3401 | percent |
| Kimi K2.5 | #61 | kimi-k2-5 | Moonshot AI | 70.2 | percent |
| GPT-5.1-Codex | #62 | gpt-5-1-codex | OpenAI | 70 | percent |
| GPT-5.1-Codex-Max | #63 | gpt-5-1-codex-max | OpenAI | 70 | percent |
| MiniMax M2.1 | #64 | minimax-m2-1 | MiniMax | 69.8639 | percent |
| GPT-OSS 120B | #65 | gpt-oss-120b | OpenAI | 69 | percent |
| MiMo-V2-Pro | #66 | mimo-v2-pro | Xiaomi | 68.8 | percent |
| Mistral Medium 3.5 128B | #67 | mistral-medium-3-5-128b | Mistral AI | 68.8 | percent |
| o4-mini | #68 | o4-mini | OpenAI | 68.7075 | percent |
| Kimi K2 Thinking | #69 | kimi-k2-thinking | Moonshot AI | 68.0952 | percent |
| GLM-4.7 | #70 | glm-4-7 | Z.AI | 67.9 | percent |
| Qwen3.6 27B | #71 | qwen3-6-27b | Alibaba Cloud | 67.6 | percent |
| Step 3.7 Flash | #72 | step-3-7-flash | StepFun | 67.3 | percent |
| MiMo-V2.5 | #73 | mimo-v2-5 | Xiaomi | 67.1429 | percent |
| GPT-OSS 20B | #74 | gpt-oss-20b | OpenAI | 65.1 | percent |
| K-Exaone | #75 | k-exaone | LG AI Research | 64.7 | percent |
| Qwen3.6-35B-A3B | #76 | qwen3-6-35b-a3b | Alibaba Cloud | 64.4 | percent |
| Claude Fable 5 | #77 | claude-fable-5 | Anthropic | 63.5 | percent |
| Nemotron 3 Nano Omni 30B A3B | #78 | nemotron-3-nano-omni-30b-a3b | NVIDIA | 63.2 | percent |
| Kimi K2.7 Code | #79 | kimi-k2-7-code | Moonshot AI | 63.1293 | percent |
| Claude Opus 4.8 | #80 | claude-opus-4-8 | Anthropic | 62.2 | percent |
| GLM-5V-Turbo | #81 | glm-5v-turbo | Z.AI | 61.1 | percent |
| DeepSeek V3.2 | #82 | deepseek-v3-2 | DeepSeek | 60.6803 | percent |
| Claude Opus 4.7 | #83 | claude-opus-4-7 | Anthropic | 58.6 | percent |
| Claude Opus 4.5 | #84 | claude-opus-4-5 | Anthropic | 58 | percent |
| Claude Opus 4.5 Thinking | #85 | claude-opus-4-5-thinking | Anthropic | 58 | percent |
| Ling 2.6 Flash | #86 | ling-2-6-flash | InclusionAI | 57.4 | percent |
| Claude Sonnet 4.5 | #87 | claude-sonnet-4-5 | Anthropic | 57.2789 | percent |
| DeepSeek V3.1 Terminus | #88 | deepseek-v3-1-terminus | DeepSeek | 57.0068 | percent |
| Ling 2.6 1T | #89 | ling-2-6-1t | InclusionAI | 56.8707 | percent |
| Trinity-Large-Preview | #90 | trinity-large-preview | Arcee AI | 56.3 | percent |
| Trinity-Large-Thinking | #91 | trinity-large-thinking | Arcee AI | 56.3 | percent |
| LFM2.5-8B-A1B | #92 | lfm2-5-8b-a1b | LiquidAI | 55.6 | percent |
| Claude Opus 4.1 | #93 | claude-opus-4-1 | Anthropic | 55.4422 | percent |
| Claude 4.1 Opus Thinking | #94 | claude-4-1-opus-thinking | Anthropic | 55.4 | percent |
| Gemini 3 Flash | #95 | gemini-3-flash | 55.1 | percent | |
| Claude Sonnet 4 | #96 | claude-sonnet-4 | Anthropic | 54.6939 | percent |
| Claude Opus 4 | #97 | claude-opus-4 | Anthropic | 53.7415 | percent |
| Grok 4 | #98 | grok-4 | xAI | 53.7 | percent |
| MiMo-V2-Omni | #99 | mimo-v2-omni | Xiaomi | 53.5 | percent |
| Claude Opus 4.6 | #100 | claude-opus-4-6 | Anthropic | 53.1 | percent |
| Grok 4.1 Fast (Reasoning) | #101 | grok-4-1-fast-reasoning | xAI | 52.7 | percent |
| Gemini 2.5 Flash | #102 | gemini-2-5-flash | 52.3129 | percent | |
| Qwen3.5 Omni Plus | #103 | qwen3-5-omni-plus | Alibaba Cloud | 51.1565 | percent |
| Grok 4 Fast | #104 | grok-4-fast | xAI | 50.5442 | percent |
| Grok 4 Fast (Reasoning) | #105 | grok-4-fast-reasoning | xAI | 50.5 | percent |
| Gemini 2.5 Pro | #106 | gemini-2-5-pro | 48.7075 | percent | |
| Claude Sonnet 3.7 | #107 | claude-sonnet-3-7 | Anthropic | 48.2993 | percent |
| Mistral Small 4 | #108 | mistral-small-4 | Mistral AI | 48.2 | percent |
| Grok 3 mini | #109 | grok-3-mini | xAI | 45.8503 | percent |
| Claude 4 Sonnet | #110 | claude-4-sonnet | Anthropic | 45.4 | percent |
| Gemma 4 E4B | #111 | gemma-4-e4b | 44.2 | percent | |
| GLM-4.6 | #112 | glm-4-6 | Z.AI | 43.4014 | percent |
| GPT-4.1 | #113 | gpt-4-1 | OpenAI | 43 | percent |
| Llama 4 Maverick | #114 | llama-4-maverick | Meta | 43 | percent |
| Kimi K2 0905 | #115 | kimi-k2-0905 | Moonshot AI | 41.7007 | percent |
| DeepSeek V3.1 | #116 | deepseek-v3-1 | DeepSeek | 41.5 | percent |
| Kimi K2 | #117 | kimi-k2 | Moonshot AI | 41.5 | percent |
| Grok Code Fast 1 | #118 | grok-code-fast-1 | xAI | 41.4 | percent |
| Claude Sonnet 4.6 | #119 | claude-sonnet-4-6 | Anthropic | 41.2 | percent |
| DeepSeek-R1 | #120 | deepseek-r1 | DeepSeek | 39.6 | percent |
| Llama 4 Scout | #121 | llama-4-scout | Meta | 39.5 | percent |
| Mistral Medium 3 | #122 | mistral-medium-3 | Mistral AI | 39.3 | percent |
| Llama 3.1 405B | #123 | llama-3-1-405b | Meta | 39 | percent |
| GPT-4.1 mini | #124 | gpt-4-1-mini | OpenAI | 38.3 | percent |
| Nemotron Ultra 253B | #125 | nemotron-ultra-253b | NVIDIA | 38.2 | percent |
| Nova Pro | #126 | nova-pro | Amazon | 38.1 | percent |
| GLM-4.5-Air | #127 | glm-4-5-air | Z.AI | 37.6 | percent |
| Grok 4.1 Fast | #128 | grok-4-1-fast | xAI | 36.5 | percent |
| Mistral Large 3 | #129 | mistral-large-3 | Mistral AI | 36.2 | percent |
| Claude 3 Haiku | #130 | claude-3-haiku | Anthropic | 36.1 | percent |
| Gemma 4 E2B | #131 | gemma-4-e2b | 36 | percent | |
| DeepSeek V3 | #132 | deepseek-v3 | DeepSeek | 34.8 | percent |
| Sarvam 105B | #133 | sarvam-105b | Sarvam | 34.4 | percent |
| GPT-4o | #134 | gpt-4o | OpenAI | 34.3 | percent |
| Solar Pro 2 | #135 | solar-pro-2 | Upstage | 33.7 | percent |
| Exaone 4.0 32B | #136 | exaone-4-0-32b | LG AI Research | 33.5 | percent |
| LFM2.5-VL-1.6B-Extract | #137 | lfm2-5-vl-1-6b-extract | LiquidAI | 33.1 | percent |
| GPT-4.1 nano | #138 | gpt-4-1-nano | OpenAI | 32 | percent |
| Gemma 3 27B | #139 | gemma-3-27b | 31.8 | percent | |
| Mistral Large 2 | #140 | mistral-large-2 | Mistral AI | 31.2 | percent |
| GPT-4o mini | #141 | gpt-4o-mini | OpenAI | 31 | percent |
| Sarvam 30B | #142 | sarvam-30b | Sarvam | 26.5 | percent |
| Granite-4.0-H-1B | #143 | granite-4-0-h-1b | IBM | 26.2 | percent |
| Exaone 4.0 1.2B | #144 | exaone-4-0-1-2b | LG AI Research | 25.3 | percent |
| Phi-4 | #145 | phi-4 | Microsoft | 23.5 | percent |
| DeepSeek R1 Distill Qwen 32B | #146 | deepseek-r1-distill-qwen-32b | DeepSeek | 22.9 | percent |
| Granite-4.0-1B | #147 | granite-4-0-1b | IBM | 20.5 | percent |
| Granite-4.0-H-350M | #148 | granite-4-0-h-350m | IBM | 17.6 | percent |
| Granite-4.0-350M | #149 | granite-4-0-350m | IBM | 16.8 | percent |
One best compatible score per canonical product · higher is better
| Rank | Model | Weights | Evidence | Score |
|---|---|---|---|---|
| 01 | MiniMax M3MiniMax | Open weights | Source carrier | 82.9% |
| 02 | Qwen3.8 MaxAlibaba Cloud | Open weights | Public reference | 82.8% |
| 03 | Nemotron 3 UltraNVIDIA | Open weights | Source carrier | 81.4% |
| 04 | Grok 4.3xAI | Closed weights | Source carrier | 81.3% |
| 05 | Qwen3.8-Flash-NextAlibaba Cloud | Open weights | Public reference | 81.3% |
| 06 | Grok 4.20xAI | Closed weights | Public reference | 81.2% |
| 07 | Qwen3.7-MaxAlibaba Cloud | Closed weights | Source carrier | 80.5% |
| 08 | Solar Open 2 250BUpstage | Open weights | Public reference | 80% |
| 09 | MiMo-V2.5-ProXiaomi | Open weights | Source carrier | 79.9% |
| 10 | Qwen3.8-27BAlibaba Cloud | Open weights | Public reference | 79.5% |
Each distinct published measurement is retained. Same-snapshot canonical and configuration projections appear once.
| Published model / configuration | Score | Evidence & protocol | Source & dates |
|---|---|---|---|
Qwen3.8-Flash-Next (xhigh default thinking configuration)Qwen3.8-Flash-NextExact identityqwen-3-8-flash-next-xhigh Canonical product: qwen-3-8-flash-next | 81.3% | Public referencesource-checkedVersion & system294 tasks system:qwen3-8-flash-next:ifbench | Qwen3.8-Flash-Next current official launch pageObserved Checked |
Qwen3.8-27B (xhigh default thinking configuration) as published in Qwen's Qwen3.8-Flash-Next comparison tableQwen3.8-27BExact identityqwen-3-8-27b-xhigh Canonical product: qwen-3-8-27b | 79.5% | Relative comparisonprovider-reportedVersion & system294 tasks system:qwen3-8-flash-next-comparison:cell:language:ifbench:1 | Qwen3.8-Flash-Next launch and official provider evaluationsObserved Checked |
DeepSeek-V4-Flash-0731 (provider-published configuration) as published in Qwen's Qwen3.8-Flash-Next comparison tableDeepSeek V4 Flash 0731Exact identitydeepseek-v4-flash-0731-deepseek-0813-release-unspecified Canonical product: deepseek-v4-flash-0731 | 79.2% | Relative comparisonprovider-reportedVersion & system294 tasks system:qwen3-8-flash-next-comparison:cell:language:ifbench:3 | Qwen3.8-Flash-Next launch and official provider evaluationsObserved Checked |
Qwen3.7-Plus (provider-published configuration) as published in Qwen's Qwen3.8-Flash-Next comparison tableQwen3.7-PlusExact identityqwen-3-7-plus-unspecified Canonical product: qwen-3-7-plus | 79.1% | Relative comparisonprovider-reportedVersion & system294 tasks system:qwen3-8-flash-next-comparison:cell:language:ifbench:2 | Qwen3.8-Flash-Next launch and official provider evaluationsObserved Checked |
Claude Opus 4.6 (Max) as published in Qwen's Qwen3.8-Flash-Next comparison tableClaude Opus 4.6Exact identityclaude-opus-4-6-max Canonical product: claude-opus-4-6 | 62.5% | Relative comparisonprovider-reportedVersion & system294 tasks system:qwen3-8-flash-next-comparison:cell:language:ifbench:4 | Qwen3.8-Flash-Next launch and official provider evaluationsObserved Checked |
DeepSeek V4 Flash max as published by UpstageDeepSeek V4 FlashExact identitydeepseek-v4-flash-max Canonical product: deepseek-v4-flash | 80.3% | Relative comparisonprovider-reportedVersion & systemstandard Solar Open 2 card English table | Solar Open 2 250B model cardObserved Checked |
Solar Open 2 250B (high)Solar Open 2 250BExact identitysolar-open2-250b-high Canonical product: solar-open2-250b | 80% | Public referenceprovider-reportedVersion & systemstandard Solar Open 2 card English table | Solar Open 2 250B model cardObserved Checked |
Qwen3.8-27B (xhigh)Qwen3.8-27BExact identityqwen-3-8-27b-xhigh Canonical product: qwen-3-8-27b | 79.5% | Public referenceprovider-reportedVersion & system294 tasks Qwen3.8-27B official text table | Qwen3.8-27B official model cardObserved Checked |
Qwen3.7-Plus as published by QwenQwen3.7-PlusExact identityqwen-3-7-plus-unspecified Canonical product: qwen-3-7-plus | 79.1% | Relative comparisonprovider-reportedVersion & system294 tasks Qwen3.8-27B official text table | Qwen3.8-27B official model cardObserved Checked |
Command A+ as published by UpstageCommand A+Exact identitycommand-a-plus-solar-open2-unspecified Canonical product: command-a-plus | 73.9% | Relative comparisonprovider-reportedVersion & systemstandard Solar Open 2 card English table | Solar Open 2 250B model cardObserved Checked |
From result to context
Instruction-following evaluation measuring constraint and format reliability.
The summary shows one best compatible source result per canonical product. Versions, effort settings and execution systems remain attached to the underlying records.
Scores from different versions or harnesses may not be interchangeable. The published source rows preserve those distinctions and their original units. A source result is not automatically an input to a current index.
This catalogue entry preserves available source evidence. Current indices admit only their specifically reviewed tracks and configurations.
Contamination risk: Unknown. Lifecycle: Active.