A task with its own rules.
Artificial Analysis knowledge and hallucination benchmark measuring factual recall across economically relevant domains.
- Organisation
- Artificial Analysis
- Version
- Current / rolling
Research / Benchmark profile
Artificial Analysis knowledge and hallucination benchmark measuring factual recall across economically relevant domains.
Observed results
20 ranked models · higher is better · labels show rank and score
| Model | Rank | Model ID | Provider | Score | Unit |
|---|---|---|---|---|---|
| Muse Spark 1.1 | #1 | muse-spark-1-1 | Meta | 40.6 | percent |
| Claude Fable 5 | #2 | claude-fable-5 | Anthropic | 40.2 | index |
| Gemini 3.1 Pro Preview | #3 | gemini-3-1-pro-preview | 32.9333 | index | |
| Claude Opus 5 | #4 | claude-opus-5 | Anthropic | 31.3 | index |
| Claude Opus 4.8 | #5 | claude-opus-4-8 | Anthropic | 27.4333 | index |
| Grok 4.5 | #6 | grok-4-5 | xAI | 26.4 | index |
| Claude Opus 4.7 | #7 | claude-opus-4-7 | Anthropic | 26.2 | index |
| Gemini 3.6 Flash | #8 | gemini-3-6-flash | 23.5 | index | |
| Gemini 3.5 Flash | #9 | gemini-3-5-flash | 22.7 | index | |
| GPT-5.6 Sol | #10 | gpt-5-6-sol | OpenAI | 21.7 | index |
| GPT-5.5 | #11 | gpt-5-5 | OpenAI | 20.1 | index |
| Kimi K3 | #12 | kimi-k3 | Moonshot AI | 18.4 | index |
| Grok 4.3 | #13 | grok-4-3 | xAI | 18.3167 | index |
| NVIDIA Nemotron 3.5 Lightning 30B-A3B | #14 | nemotron-3-5-lightning-30b-a3b | NVIDIA | 17.5 | index |
| Grok 4.20 | #15 | grok-4-20 | xAI | 15.35 | index |
| Claude Sonnet 5 | #16 | claude-sonnet-5 | Anthropic | 15.3167 | index |
| Qwen3.7-Max | #17 | qwen-3-7-max | Alibaba Cloud | 14.1 | index |
| Claude Opus 4.6 | #18 | claude-opus-4-6 | Anthropic | 13.5 | index |
| Claude Opus 4.5 Thinking | #19 | claude-opus-4-5-thinking | Anthropic | 13.3 | index |
| Claude Opus 4.5 | #20 | claude-opus-4-5 | Anthropic | 13.2667 | index |
| Claude Sonnet 4.6 | #21 | claude-sonnet-4-6 | Anthropic | 12.3667 | index |
| Gemini 3 Flash | #22 | gemini-3-flash | 11.5667 | index | |
| Qwen 3.6 Max (preview) | #23 | qwen3-6-max-preview | Alibaba Cloud | 10.2 | index |
| Qwen3.6 Max | #24 | qwen3-6-max | Alibaba Cloud | 10.2 | index |
| GPT-5.3-Codex | #25 | gpt-5-3-codex | OpenAI | 9.9 | index |
| Gemini 3.5 Flash-Lite | #26 | gemini-3-5-flash-lite | 6.9 | index | |
| Kimi K2.6 | #27 | kimi-k2-6 | Moonshot AI | 6.4167 | index |
| GPT-5.4 | #28 | gpt-5-4 | OpenAI | 5.7 | index |
| GPT-5.1 | #29 | gpt-5-1 | OpenAI | 5.6 | index |
| MiMo-V2-Pro | #30 | mimo-v2-pro | Xiaomi | 4.9167 | index |
| Muse Spark | #31 | muse-spark | Meta | 4.1 | index |
| GLM-5.2 | #32 | glm-5-2 | Z.AI | 4 | index |
| Grok 4 | #33 | grok-4 | xAI | 3.8 | index |
| MiMo-V2.5-Pro | #34 | mimo-v2-5-pro | Xiaomi | 3.6 | index |
| Qwen3.6 Plus | #35 | qwen3-6-plus | Alibaba Cloud | 2.7 | index |
| Qwen3.7-Plus | #36 | qwen-3-7-plus | Alibaba Cloud | 2.4 | index |
| Inkling | #37 | inkling | Thinking Machines Lab | 2.1 | index |
| GLM-5 | #38 | glm-5 | Z.AI | 2 | index |
| GLM-5.1 | #39 | glm-5-1 | Z.AI | 1.9333 | index |
| MiniMax M3 | #40 | minimax-m3 | MiniMax | 1.4 | index |
| MiniMax M2.7 | #41 | minimax-m2-7 | MiniMax | 0.7 | index |
| Claude Sonnet 4.5 | #42 | claude-sonnet-4-5 | Anthropic | 0.3333 | index |
| Claude Sonnet 3.7 | #43 | claude-sonnet-3-7 | Anthropic | 0.3167 | index |
| Claude Sonnet 4 | #44 | claude-sonnet-4 | Anthropic | 0.1833 | index |
| GPT-5.6 Terra | #45 | gpt-5-6-terra | OpenAI | -0.2 | index |
| Nemotron 3 Ultra | #46 | nemotron-3-ultra | NVIDIA | -0.8 | index |
| GPT-5.2 | #47 | gpt-5-2 | OpenAI | -1 | index |
| GPT-5.2 Codex | #48 | gpt-5-2-codex | OpenAI | -2.4833 | index |
| Command A+ | #49 | command-a-plus | Cohere | -3.9833 | index |
| Claude Haiku 4.5 | #50 | claude-haiku-4-5 | Anthropic | -4.2167 | index |
| GPT-5.1-Codex | #51 | gpt-5-1-codex | OpenAI | -6 | index |
| GPT-5.1-Codex-Max | #52 | gpt-5-1-codex-max | OpenAI | -6 | index |
| Grok 3 mini | #53 | grok-3-mini | xAI | -6.05 | index |
| GPT-5 Codex | #54 | gpt-5-codex | OpenAI | -6.7667 | index |
| GPT-5 | #55 | gpt-5 | OpenAI | -8.0833 | index |
| Kimi K2.5 | #56 | kimi-k2-5 | Moonshot AI | -8.1 | index |
| Claude 4 Sonnet | #57 | claude-4-sonnet | Anthropic | -9.2 | index |
| MiMo-V2.5 | #58 | mimo-v2-5 | Xiaomi | -9.3333 | index |
| DeepSeek V4 Pro | #59 | deepseek-v4-pro | DeepSeek | -9.7 | index |
| o1 | #60 | o1 | OpenAI | -10.5 | index |
| GPT-4o | #61 | gpt-4o | OpenAI | -10.7 | index |
| Kimi K2.7 Code | #62 | kimi-k2-7-code | Moonshot AI | -10.7 | index |
| GPT-5 mini | #63 | gpt-5-mini | OpenAI | -10.7833 | index |
| GPT-5.6 Luna | #64 | gpt-5-6-luna | OpenAI | -11.2 | index |
| Qwen3.5 Omni Plus | #65 | qwen3-5-omni-plus | Alibaba Cloud | -12.3 | index |
| Gemini 2.5 Pro | #66 | gemini-2-5-pro | -14.3 | index | |
| GLM-5-Turbo | #67 | glm-5-turbo | Z.AI | -15.0833 | index |
| o3 | #68 | o3 | OpenAI | -15.2667 | index |
| Gemini 3.1 Flash-Lite | #69 | gemini-3-1-flash-lite | -15.5 | index | |
| Llama 3.1 405B | #70 | llama-3-1-405b | Meta | -17.3 | index |
| MiMo-V2-Omni | #71 | mimo-v2-omni | Xiaomi | -17.4 | index |
| Hy3 | #72 | hy3 | Tencent | -18.5 | index |
| Hy3 Preview | #73 | hy3-preview | Tencent | -18.5 | index |
| GPT-5.4 mini | #74 | gpt-5-4-mini | OpenAI | -18.6833 | index |
| GLM-5V-Turbo | #75 | glm-5v-turbo | Z.AI | -18.9833 | index |
| Qwen3.6 27B | #76 | qwen3-6-27b | Alibaba Cloud | -19.7833 | index |
| Gemma 4 E4B | #77 | gemma-4-e4b | -20 | index | |
| Kimi K2 Thinking | #78 | kimi-k2-thinking | Moonshot AI | -20.4833 | index |
| DeepSeek V3.2 | #79 | deepseek-v3-2 | DeepSeek | -20.8833 | index |
| Qwen3.6-35B-A3B | #80 | qwen3-6-35b-a3b | Alibaba Cloud | -21.4 | index |
| DeepSeek V4 Flash | #81 | deepseek-v4-flash | DeepSeek | -22.3 | index |
| Gemma 4 E2B | #82 | gemma-4-e2b | -24 | index | |
| DeepSeek V3.1 Terminus | #83 | deepseek-v3-1-terminus | DeepSeek | -24.3167 | index |
| Kimi K2 0905 | #84 | kimi-k2-0905 | Moonshot AI | -26.4667 | index |
| DeepSeek-R1 | #85 | deepseek-r1 | DeepSeek | -27.1 | index |
| Kimi K2 | #86 | kimi-k2 | Moonshot AI | -27.5 | index |
| DeepSeek V3.1 | #87 | deepseek-v3-1 | DeepSeek | -28.4 | index |
| Grok 4 Fast | #88 | grok-4-fast | xAI | -28.4 | index |
| Grok 4 Fast (Reasoning) | #89 | grok-4-fast-reasoning | xAI | -28.4 | index |
| Grok 4.1 Fast (Reasoning) | #90 | grok-4-1-fast-reasoning | xAI | -28.7 | index |
| GPT-5.4 nano | #91 | gpt-5-4-nano | OpenAI | -29.5 | index |
| Qwen3.5 397B A17B | #92 | qwen3-5-397b | Alibaba Cloud | -29.7833 | index |
| Mistral Small 4 | #93 | mistral-small-4 | Mistral AI | -29.9 | index |
| Mistral Medium 3 | #94 | mistral-medium-3 | Mistral AI | -31.5 | index |
| GLM-4.6 | #95 | glm-4-6 | Z.AI | -31.6 | index |
| MiniMax M2.1 | #96 | minimax-m2-1 | MiniMax | -32.8333 | index |
| LFM2.5-8B-A1B | #97 | lfm2-5-8b-a1b | LiquidAI | -33.3 | index |
| Mistral Large 2 | #98 | mistral-large-2 | Mistral AI | -34 | index |
| Qwen3 Max | #99 | qwen3-max | Alibaba Cloud | -34.4167 | index |
| GLM-4.7 | #100 | glm-4-7 | Z.AI | -34.6 | index |
| Gemini 2.5 Flash | #101 | gemini-2-5-flash | -34.7167 | index | |
| o4-mini | #102 | o4-mini | OpenAI | -35.75 | index |
| Grok Code Fast 1 | #103 | grok-code-fast-1 | xAI | -36 | index |
| GPT-4.1 | #104 | gpt-4-1 | OpenAI | -36.2 | index |
| Mistral Medium 3.5 128B | #105 | mistral-medium-3-5-128b | Mistral AI | -36.3 | index |
| Mistral Medium 3.5 | #106 | mistral-medium-3-5 | Mistral AI | -36.3167 | index |
| Step 3.7 Flash | #107 | step-3-7-flash | StepFun | -37.5 | index |
| Mistral Large 3 | #108 | mistral-large-3 | Mistral AI | -39.4 | index |
| Qwen3.5 122B | #109 | qwen3-5-122b | Alibaba Cloud | -39.5833 | index |
| Qwen3.5-122B-A10B | #110 | qwen3-5-122b-a10b | Alibaba Cloud | -39.6 | index |
| MiniMax M2.5 | #111 | minimax-m2-5 | MiniMax | -39.7 | index |
| DeepSeek V3 | #112 | deepseek-v3 | DeepSeek | -41.3 | index |
| Llama 4 Maverick | #113 | llama-4-maverick | Meta | -41.8 | index |
| Qwen3.5 27B | #114 | qwen3-5-27b | Alibaba Cloud | -42 | index |
| Nemotron 3 Super | #115 | nemotron-3-super | NVIDIA | -42.0667 | index |
| Trinity-Large-Preview | #116 | trinity-large-preview | Arcee AI | -44.2 | index |
| Trinity-Large-Thinking | #117 | trinity-large-thinking | Arcee AI | -44.2 | index |
| Gemma 4 31B | #118 | gemma-4-31b | -45.4 | index | |
| Nemotron Ultra 253B | #119 | nemotron-ultra-253b | NVIDIA | -45.5 | index |
| Qwen3.5 35B | #120 | qwen3-5-35b | Alibaba Cloud | -46.3833 | index |
| Qwen3.5-35B-A3B | #121 | qwen3-5-35b-a3b | Alibaba Cloud | -46.4 | index |
| MiniMax M2 | #122 | minimax-m2 | MiniMax | -46.9333 | index |
| Claude 3 Haiku | #123 | claude-3-haiku | Anthropic | -47.6 | index |
| Nova Pro | #124 | nova-pro | Amazon | -47.6 | index |
| Nova 2 Pro | #125 | nova-2-pro | Amazon | -48.05 | index |
| Gemma 4 26B | #126 | gemma-4-26b | -48.0667 | index | |
| Gemma 4 26B A4B | #127 | gemma-4-26b-a4b | -48.1 | index | |
| GPT-OSS 120B | #128 | gpt-oss-120b | OpenAI | -50 | index |
| GPT-4.1 mini | #129 | gpt-4-1-mini | OpenAI | -50.1 | index |
| Grok 4.1 Fast | #130 | grok-4-1-fast | xAI | -50.9 | index |
| Ling 2.6 1T | #131 | ling-2-6-1t | InclusionAI | -51 | index |
| Nemotron 3 Nano 30B | #132 | nemotron-3-nano-30b | NVIDIA | -51.6 | index |
| Gemma 4 12B Unified | #133 | gemma-4-12b | -51.8833 | index | |
| Llama 4 Scout | #134 | llama-4-scout | Meta | -52.4 | index |
| Nemotron 3 Nano Omni 30B A3B | #135 | nemotron-3-nano-omni-30b-a3b | NVIDIA | -56 | index |
| GPT-4.1 nano | #136 | gpt-4-1-nano | OpenAI | -56.4 | index |
| Phi-4 | #137 | phi-4 | Microsoft | -56.7 | index |
| K-Exaone | #138 | k-exaone | LG AI Research | -57.9 | index |
| Sarvam 105B | #139 | sarvam-105b | Sarvam | -59.5 | index |
| Solar Pro 2 | #140 | solar-pro-2 | Upstage | -61.7 | index |
| Exaone 4.0 32B | #141 | exaone-4-0-32b | LG AI Research | -62.3 | index |
| GLM-4.5-Air | #142 | glm-4-5-air | Z.AI | -62.5 | index |
| GPT-OSS 20B | #143 | gpt-oss-20b | OpenAI | -63.9 | index |
| Ling 2.6 Flash | #144 | ling-2-6-flash | InclusionAI | -65.7 | index |
| Gemma 3 27B | #145 | gemma-3-27b | -65.9 | index | |
| Sarvam 30B | #146 | sarvam-30b | Sarvam | -72 | index |
| Granite-4.0-350M | #147 | granite-4-0-350m | IBM | -72.1 | index |
| Granite-4.0-H-1B | #148 | granite-4-0-h-1b | IBM | -73.6 | index |
| Granite-4.0-1B | #149 | granite-4-0-1b | IBM | -81.8 | index |
| Exaone 4.0 1.2B | #150 | exaone-4-0-1-2b | LG AI Research | -82.6 | index |
| LFM2.5-VL-1.6B-Extract | #151 | lfm2-5-vl-1-6b-extract | LiquidAI | -83.9 | index |
| Granite-4.0-H-350M | #152 | granite-4-0-h-350m | IBM | -87.2 | index |
One best compatible score per canonical product · higher is better
| Rank | Model | Weights | Evidence | Score |
|---|---|---|---|---|
| 01 | Muse Spark 1.1Meta | Closed weights | Public reference | 40.6% |
| 02 | Claude Fable 5Anthropic | Closed weights | Public reference | 40.2 index |
| 03 | Gemini 3.1 Pro PreviewGoogle | Closed weights | Public reference | 32.9 index |
| 04 | Claude Opus 5Anthropic | Closed weights | Public reference | 31.3 index |
| 05 | Claude Opus 4.8Anthropic | Closed weights | Public reference | 27.4 index |
| 06 | Grok 4.5xAI | Closed weights | Public reference | 26.4 index |
| 07 | Claude Opus 4.7Anthropic | Closed weights | Public reference | 26.2 index |
| 08 | Gemini 3.6 FlashGoogle | Closed weights | Public reference | 23.5 index |
| 09 | Gemini 3.5 FlashGoogle | Closed weights | Public reference | 22.7 index |
| 10 | GPT-5.6 SolOpenAI | Closed weights | Public reference | 21.7 index |
Each distinct published measurement is retained. Same-snapshot canonical and configuration projections appear once.
| Published model / configuration | Score | Evidence & protocol | Source & dates |
|---|---|---|---|
NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16; NVIDIA release evaluation; temperature=1.0; top_p=0.95NVIDIA Nemotron 3.5 Lightning 30B-A3BExact identitynemotron-3-5-lightning-30b-a3b-default Canonical product: nemotron-3-5-lightning-30b-a3b | 17.5 index | Public referenceprovider-reportedVersion & systemCurrent NVIDIA NeMo Gym / NeMo Evaluator SDK consistent release harness | NVIDIA Nemotron 3.5 Lightning 30B-A3B BF16 model cardObserved Checked |
Exact BenchLM registry variant Claude Fable 5; bulk export does not retain a complete upstream harness configuration.Claude Fable 5Exact identitySource label without a registered configuration ID Canonical product: claude-fable-5 | 40.2 index | Public referencesource-checkedVersion & system2026 Source-native system | BenchLM public datasets — 2026-08-01Observed Checked |
Exact BenchLM registry variant Claude Opus 5; bulk export does not retain a complete upstream harness configuration.Claude Opus 5Exact identitySource label without a registered configuration ID Canonical product: claude-opus-5 | 31.3 index | Public referencesource-checkedVersion & system2026 Source-native system | BenchLM public datasets — 2026-08-01Observed Checked |
Exact BenchLM registry variant Claude Opus 4.8; bulk export does not retain a complete upstream harness configuration.Claude Opus 4.8Exact identitySource label without a registered configuration ID Canonical product: claude-opus-4-8 | 27.4 index | Public referencesource-checkedVersion & system2026 Source-native system | BenchLM public datasets — 2026-08-01Observed Checked |
Exact BenchLM registry variant Grok 4.5; bulk export does not retain a complete upstream harness configuration.Grok 4.5Exact identitySource label without a registered configuration ID Canonical product: grok-4-5 | 26.4 index | Public referencesource-checkedVersion & system2026 Source-native system | BenchLM public datasets — 2026-08-01Observed Checked |
Exact BenchLM registry variant Claude Opus 4.7 (Adaptive); bulk export does not retain a complete upstream harness configuration.Claude Opus 4.7Exact identityclaude-opus-4-7-max Canonical product: claude-opus-4-7 | 26.2 index | Public referencesource-checkedVersion & system2026 Source-native system | BenchLM public datasets — 2026-08-01Observed Checked |
Exact BenchLM registry variant Gemini 3.6 Flash; bulk export does not retain a complete upstream harness configuration.Gemini 3.6 FlashExact identitySource label without a registered configuration ID Canonical product: gemini-3-6-flash | 23.5 index | Public referencesource-checkedVersion & system2026 Source-native system | BenchLM public datasets — 2026-08-01Observed Checked |
Exact BenchLM registry variant Gemini 3.5 Flash; bulk export does not retain a complete upstream harness configuration.Gemini 3.5 FlashExact identitySource label without a registered configuration ID Canonical product: gemini-3-5-flash | 22.7 index | Public referencesource-checkedVersion & system2026 Source-native system | BenchLM public datasets — 2026-08-01Observed Checked |
Exact BenchLM registry variant GPT-5.6 Sol; bulk export does not retain a complete upstream harness configuration.GPT-5.6 SolExact identitySource label without a registered configuration ID Canonical product: gpt-5-6-sol | 21.7 index | Public referencesource-checkedVersion & system2026 Source-native system | BenchLM public datasets — 2026-08-01Observed Checked |
Exact BenchLM registry variant GPT-5.5; bulk export does not retain a complete upstream harness configuration.GPT-5.5Exact identitySource label without a registered configuration ID Canonical product: gpt-5-5 | 20.1 index | Public referencesource-checkedVersion & system2026 Source-native system | BenchLM public datasets — 2026-08-01Observed Checked |
| Model | Result | Execution & attribution | Original source |
|---|---|---|---|
| Qwen3.8 Max (0902)omniscience · exact AA profile b112af07-3bd5-4647-b09f-b23204361bb2 | 11.983native score | AA current published profile; effort label not exportedindependent evaluatorRetained as additional evidence. Existing Core point is unchanged pending an accepted complete own-evidence amendment; incompatible versions remain separate. | Artificial Analysis ↗Publication date not recorded · checked 2026-09-16 |
From result to context
Artificial Analysis knowledge and hallucination benchmark measuring factual recall across economically relevant domains.
The summary shows one best compatible source result per canonical product. Versions, effort settings and execution systems remain attached to the underlying records.
Scores from different versions or harnesses may not be interchangeable. The published source rows preserve those distinctions and their original units. A source result is not automatically an input to a current index.
A reviewed track from this benchmark is represented in the following current indices. Each index applies its own protocol and eligibility rules.
Contamination risk: Unknown. Lifecycle: Active.