A task with its own rules.
Factual short-answer evaluation derived from SimpleQA with stricter verification protocols against known ground-truth answers.
- Organisation
- OpenAI / Google
- Version
- verified
Research / Benchmark profile
Factual short-answer evaluation derived from SimpleQA with stricter verification protocols against known ground-truth answers.
Observed results
5 ranked models · higher is better · labels show rank and score
| Model | Rank | Model ID | Provider | Score | Unit |
|---|---|---|---|---|---|
| Gemini 3.7 Flash | #1 | gemini-3-7-flash | 69.2 | percent | |
| Muse Spark 1.2 | #2 | muse-spark-1-2 | Meta | 60.301508 | percent |
| Grok 4.6 | #3 | grok-4-6 | xAI | 48.9 | percent |
| Qwen3.8 Max | #4 | qwen-3-8-max | Alibaba Cloud | 45.8 | percent |
| GLM-5.3 | #5 | glm-5-3 | Z.AI | 41 | percent |
One best compatible score per canonical product · higher is better
| Rank | Model | Weights | Evidence | Score |
|---|---|---|---|---|
| 01 | Gemini 3.7 FlashGoogle | Closed weights | Public reference | 69.2% |
| 02 | Muse Spark 1.2Meta | Open announced | Public reference | 60.3% |
| 03 | Grok 4.6xAI | Closed weights | Public reference | 48.9% |
| 04 | Qwen3.8 MaxAlibaba Cloud | Open weights | Public reference | 45.8% |
| 05 | GLM-5.3Z.AI | Open announced | Public reference | 41% |
Each distinct published measurement is retained. Same-snapshot canonical and configuration projections appear once.
| Published model / configuration | Score | Evidence & protocol | Source & dates |
|---|---|---|---|
GLM-5.3 (max)GLM-5.3Exact identityglm-5-3-max Canonical product: glm-5-3 | 41% | Public referenceindependently-verifiedVersion & system1.2.0 Epoch AI benchmark runner | Epoch AI permanent refresh sourceObserved Checked |
Gemini 3.7 Flash (high)Gemini 3.7 FlashExact identitygemini-3-7-flash-high Canonical product: gemini-3-7-flash | 69.2% | Public referenceindependently-verifiedVersion & system1.2.0 Epoch AI benchmark runner | Epoch AI permanent refresh sourceObserved Checked |
Muse Spark 1.2 (xhigh)Muse Spark 1.2Exact identitymuse-spark-1-2-xhigh Canonical product: muse-spark-1-2 | 60.3% | Public referenceindependently-verifiedVersion & system1.2.0 Epoch AI benchmark runner | Epoch AI permanent refresh sourceObserved Checked |
Grok 4.6 (xhigh)Grok 4.6Exact identitygrok-4-6-xhigh Canonical product: grok-4-6 | 48.9% | Public referenceindependently-verifiedVersion & system1.2.0 Epoch AI benchmark runner | Epoch AI permanent refresh sourceObserved Checked |
Qwen3.8 Max (xhigh)Qwen3.8 MaxExact identityqwen-3-8-max-xhigh Canonical product: qwen-3-8-max | 45.8% | Public referenceindependently-verifiedVersion & system1.2.0 Epoch AI benchmark runner | Epoch AI permanent refresh sourceObserved Checked |
From result to context
Factual short-answer evaluation derived from SimpleQA with stricter verification protocols against known ground-truth answers.
The summary shows one best compatible source result per canonical product. Versions, effort settings and execution systems remain attached to the underlying records.
Scores from different versions or harnesses may not be interchangeable. The published source rows preserve those distinctions and their original units. A source result is not automatically an input to a current index.
This catalogue entry preserves available source evidence. Current indices admit only their specifically reviewed tracks and configurations.
Contamination risk: Medium. Lifecycle: Archived.