Reasoning
Mathematical and scientific problem solving.
Model profile / DeepSeek
DeepSeek's experimental multimodal V4 Flash model for text, image understanding and agent workflows.
Four areas of evidence
Mathematical and scientific problem solving.
Software engineering and programming.
Factual reliability and instruction following.
Tool use, workflows and professional tasks.
Current index snapshot · 2026-09-16. Each category uses its own scale; scores across categories are not directly comparable. Counts describe published catalogue evidence. A point lead does not establish superiority.
Performance in context
DeepSeek V4 Flash Vision Exp has no current Capabilities rating. Explore the rated models below.
Lumina Capabilities Index
The current Capabilities cohort, with available operational measurements.
Chart loads as you explore
Lumina Capabilities Index
The current Capabilities cohort, with available operational measurements.
Chart loads as you explore
Explore the detail
No current index estimate. Available benchmark measurements and source records are retained below.
| Benchmark | Result | Published configuration | Source |
|---|---|---|---|
| AA Long Context Reasoningstandard | 78% | DeepSeek V4 Flash Vision Exp (max; Artificial Analysis independent run)Artificial Analysis AA-LCR evaluation | DeepSeek V4 Flash Vision Exp individual evaluations ↗Observed 2026-08-27 · checked 2026-08-27independently-verified · reference-only |
| CritPtstandard | 10.857% | DeepSeek V4 Flash Vision Exp (max; Artificial Analysis independent run)Artificial Analysis CritPt evaluation | DeepSeek V4 Flash Vision Exp individual evaluations ↗Observed 2026-08-27 · checked 2026-08-27independently-verified · reference-only |
| GDPval-AA v2v2 | 58.141% | DeepSeek V4 Flash Vision Exp (max; Artificial Analysis independent run)Artificial Analysis GDPval-AA v2 | DeepSeek V4 Flash Vision Exp individual evaluations ↗Observed 2026-08-27 · checked 2026-08-27independently-verified · reference-only |
| GPQA Diamonddiamond | 91.313% | DeepSeek V4 Flash Vision Exp (max; Artificial Analysis independent run)Artificial Analysis GPQA Diamond evaluation | DeepSeek V4 Flash Vision Exp individual evaluations ↗Observed 2026-08-27 · checked 2026-08-27independently-verified · reference-only |
| Humanity's Last Examtext-only current | 34.476% | DeepSeek V4 Flash Vision Exp (max; Artificial Analysis independent run)Artificial Analysis text-only HLE evaluation | DeepSeek V4 Flash Vision Exp individual evaluations ↗Observed 2026-08-27 · checked 2026-08-27independently-verified · reference-only |
| SciCode2024 | 46.644% | DeepSeek V4 Flash Vision Exp (max; Artificial Analysis independent run)Artificial Analysis SciCode evaluation | DeepSeek V4 Flash Vision Exp individual evaluations ↗Observed 2026-08-27 · checked 2026-08-27independently-verified · reference-only |
| τ³-Banking3 | 41.031% | DeepSeek V4 Flash Vision Exp (max; Artificial Analysis independent run)Artificial Analysis tau3-Banking evaluation | DeepSeek V4 Flash Vision Exp individual evaluations ↗Observed 2026-08-27 · checked 2026-08-27independently-verified · reference-only |
| Terminal-Bench2.1 | 74.157% | DeepSeek V4 Flash Vision Exp (max; Artificial Analysis independent run)Artificial Analysis Terminal-Bench v2.1 in e2b | DeepSeek V4 Flash Vision Exp individual evaluations ↗Observed 2026-08-27 · checked 2026-08-27independently-verified · reference-only |
| Agents Last Exam2026 | 27.3% | Provider chart does not specify an exact evaluation systemSource-native system | DeepSeek-V4-Flash-Vision-Exp release and provider evaluation ↗Observed 2026-08-21 · checked 2026-08-24provider-reported · reference-only |
| AutomationBench2026 public set | 25.7% | DeepSeek Harness Minimal Mode; max effort; top_p=0.95; temperature=1.0DeepSeek Harness Minimal Mode | DeepSeek-V4-Flash-Vision-Exp release and provider evaluation ↗Observed 2026-08-21 · checked 2026-08-24provider-reported · reference-only |
No specialist task observations retained for this model.