A task with its own rules.
An evaluation of whether language models follow verifiable natural-language instructions.
- Organisation
- Google Research
- Version
- Current / rolling
Instruction Following / Benchmark profile
An evaluation of whether language models follow verifiable natural-language instructions.
Observed results
20 ranked models · higher is better · labels show rank and score
| Model | Rank | Model ID | Provider | Score | Unit |
|---|---|---|---|---|---|
| Qwen3.7-Plus | #1 | qwen-3-7-plus | Alibaba Cloud | 94.6 | percent |
| Qwen3.6 Plus | #2 | qwen3-6-plus | Alibaba Cloud | 94.3 | percent |
| Qwen3.7-Max | #3 | qwen-3-7-max | Alibaba Cloud | 94.3 | percent |
| Kimi K2.5 | #4 | kimi-k2-5 | Moonshot AI | 93.9 | percent |
| GLM-5 | #5 | glm-5 | Z.AI | 92.6 | percent |
| Claude Opus 4.5 | #6 | claude-opus-4-5 | Anthropic | 90.9 | percent |
| LongCat-2.0 | #7 | longcat-2-0 | Meituan | 90 | percent |
| MiniMax M3 | #8 | minimax-m3 | MiniMax | 82.9 | percent |
| Grok 4.3 | #9 | grok-4-3 | xAI | 81.3 | percent |
| MiMo-V2.5-Pro | #10 | mimo-v2-5-pro | Xiaomi | 79.9 | percent |
| Gemini 3.1 Flash-Lite | #11 | gemini-3-1-flash-lite | 77.2 | percent | |
| Gemini 3.1 Pro Preview | #12 | gemini-3-1-pro-preview | 77.1 | percent | |
| Gemini 3.5 Flash | #13 | gemini-3-5-flash | 76.3 | percent | |
| GLM-5.1 | #14 | glm-5-1 | Z.AI | 76.3 | percent |
| GPT-5.4 nano | #15 | gpt-5-4-nano | OpenAI | 75.9 | percent |
| GPT-5.5 | #16 | gpt-5-5 | OpenAI | 75.9 | percent |
| Muse Spark | #17 | muse-spark | Meta | 75.9 | percent |
| MiniMax M2.7 | #18 | minimax-m2-7 | MiniMax | 75.7 | percent |
| Gemma 4 31B | #19 | gemma-4-31b | 75.6 | percent | |
| GPT-5.3-Codex | #20 | gpt-5-3-codex | OpenAI | 75.4 | percent |
| GPT-5.4 | #21 | gpt-5-4 | OpenAI | 73.9 | percent |
| GLM-5.2 | #22 | glm-5-2 | Z.AI | 73.3 | percent |
| GLM-5-Turbo | #23 | glm-5-turbo | Z.AI | 73.2 | percent |
| GPT-5.1 | #24 | gpt-5-1 | OpenAI | 72.9 | percent |
| GPT-5.6 Sol | #25 | gpt-5-6-sol | OpenAI | 72.7 | percent |
| GPT-5.6 Terra | #26 | gpt-5-6-terra | OpenAI | 71.2 | percent |
| MiMo-V2-Pro | #27 | mimo-v2-pro | Xiaomi | 68.8 | percent |
| Claude Fable 5 | #28 | claude-fable-5 | Anthropic | 63.5 | percent |
| Claude Opus 4.8 | #29 | claude-opus-4-8 | Anthropic | 62.2 | percent |
| GLM-5V-Turbo | #30 | glm-5v-turbo | Z.AI | 61.1 | percent |
| Gemini 3 Flash | #31 | gemini-3-flash | 55.1 | percent | |
| MiMo-V2-Omni | #32 | mimo-v2-omni | Xiaomi | 53.5 | percent |
| Mistral Small 4 | #33 | mistral-small-4 | Mistral AI | 48.2 | percent |
| Claude Opus 4.6 | #34 | claude-opus-4-6 | Anthropic | 44.6 | percent |
| Claude Opus 4.7 | #35 | claude-opus-4-7 | Anthropic | 43.6 | percent |
| Claude Sonnet 4.6 | #36 | claude-sonnet-4-6 | Anthropic | 41.2 | percent |
| Mistral Large 3 | #37 | mistral-large-3 | Mistral AI | 36.2 | percent |
One best compatible score per canonical product · higher is better
| Rank | Model | Weights | Evidence | Score |
|---|---|---|---|---|
| 01 | Qwen3.7-PlusAlibaba Cloud | Closed weights | Source carrier | 94.6% |
| 02 | Qwen3.6 PlusAlibaba Cloud | Closed weights | Source carrier | 94.3% |
| 03 | Qwen3.7-MaxAlibaba Cloud | Closed weights | Source carrier | 94.3% |
| 04 | Kimi K2.5Moonshot AI | Open weights | Source carrier | 93.9% |
| 05 | GLM-5Z.AI | Open weights | Source carrier | 92.6% |
| 06 | Claude Opus 4.5Anthropic | Closed weights | Source carrier | 90.9% |
| 07 | LongCat-2.0Meituan | Open weights | Public reference | 90% |
| 08 | MiniMax M3MiniMax | Open weights | Source carrier | 82.9% |
| 09 | Grok 4.3xAI | Closed weights | Source carrier | 81.3% |
| 10 | MiMo-V2.5-ProXiaomi | Open weights | Source carrier | 79.9% |
Each distinct published measurement is retained. Same-snapshot canonical and configuration projections appear once.
| Published model / configuration | Score | Evidence & protocol | Source & dates |
|---|---|---|---|
Exact model variant as listed on the BenchLM public benchmark leaderboard page.Qwen3.7-PlusExact identitySource label without a registered configuration ID Canonical product: qwen-3-7-plus | 94.6% | Source carriersource-checkedVersion & systemCurrent BenchLM aggregated public evaluation | BenchLM public benchmark leaderboardsObserved Checked |
Exact model variant as listed on the BenchLM public benchmark leaderboard page.Qwen3.7-MaxExact identitySource label without a registered configuration ID Canonical product: qwen-3-7-max | 94.3% | Source carriersource-checkedVersion & systemCurrent BenchLM aggregated public evaluation | BenchLM public benchmark leaderboardsObserved Checked |
Exact model variant as listed on the BenchLM public benchmark leaderboard page.Qwen3.6 PlusExact identitySource label without a registered configuration ID Canonical product: qwen3-6-plus | 94.3% | Source carriersource-checkedVersion & systemCurrent BenchLM aggregated public evaluation | BenchLM public benchmark leaderboardsObserved Checked |
Exact model variant as listed on the BenchLM public benchmark leaderboard page.Kimi K2.5Exact identitySource label without a registered configuration ID Canonical product: kimi-k2-5 | 93.9% | Source carriersource-checkedVersion & systemCurrent BenchLM aggregated public evaluation | BenchLM public benchmark leaderboardsObserved Checked |
Exact model variant as listed on the BenchLM public benchmark leaderboard page.GLM-5Exact identitySource label without a registered configuration ID Canonical product: glm-5 | 92.6% | Source carriersource-checkedVersion & systemCurrent BenchLM aggregated public evaluation | BenchLM public benchmark leaderboardsObserved Checked |
Exact model variant as listed on the BenchLM public benchmark leaderboard page.Claude Opus 4.5Exact identitySource label without a registered configuration ID Canonical product: claude-opus-4-5 | 90.9% | Source carriersource-checkedVersion & systemCurrent BenchLM aggregated public evaluation | BenchLM public benchmark leaderboardsObserved Checked |
Exact model variant as listed on the BenchLM public benchmark leaderboard page.MiniMax M3Exact identitySource label without a registered configuration ID Canonical product: minimax-m3 | 82.9% | Source carriersource-checkedVersion & systemCurrent BenchLM aggregated public evaluation | BenchLM public benchmark leaderboardsObserved Checked |
Exact model variant as listed on the BenchLM public benchmark leaderboard page.Grok 4.3Exact identitySource label without a registered configuration ID Canonical product: grok-4-3 | 81.3% | Source carriersource-checkedVersion & systemCurrent BenchLM aggregated public evaluation | BenchLM public benchmark leaderboardsObserved Checked |
Exact model variant as listed on the BenchLM public benchmark leaderboard page.MiMo-V2.5-ProExact identitySource label without a registered configuration ID Canonical product: mimo-v2-5-pro | 79.9% | Source carriersource-checkedVersion & systemCurrent BenchLM aggregated public evaluation | BenchLM public benchmark leaderboardsObserved Checked |
Exact model variant as listed on the BenchLM public benchmark leaderboard page.Gemini 3.1 Flash-LiteExact identitySource label without a registered configuration ID Canonical product: gemini-3-1-flash-lite | 77.2% | Source carriersource-checkedVersion & systemCurrent BenchLM aggregated public evaluation | BenchLM public benchmark leaderboardsObserved Checked |
From result to context
An evaluation of whether language models follow verifiable natural-language instructions.
The summary shows one best compatible source result per canonical product. Versions, effort settings and execution systems remain attached to the underlying records.
Scores from different versions or harnesses may not be interchangeable. The published source rows preserve those distinctions and their original units. A source result is not automatically an input to a current index.
This catalogue entry preserves available source evidence. Current indices admit only their specifically reviewed tracks and configurations.
Contamination risk: Unknown. Lifecycle: Active.