A task with its own rules.
A visual counting benchmark that tests whether a model can count objects and entities reliably in complex scenes.
- Organisation
- Qwen
- Version
- 2026
Multimodal / Benchmark profile
A visual counting benchmark that tests whether a model can count objects and entities reliably in complex scenes.
Observed results
2 ranked models · higher is better · labels show rank and score
| Model | Rank | Model ID | Provider | Score | Unit |
|---|---|---|---|---|---|
| Qwen3.6 27B | #1 | qwen3-6-27b | Alibaba Cloud | 97.8 | percent |
| LFM2.5-VL-450M | #2 | lfm2-5-vl-450m | LiquidAI | 73.31 | percent |
One best compatible score per canonical product · higher is better
| Rank | Model | Weights | Evidence | Score |
|---|---|---|---|---|
| 01 | Qwen3.6 27BAlibaba Cloud | Open weights | Source carrier | 97.8% |
| 02 | LFM2.5-VL-450MLiquidAI | Open weights | Source carrier | 73.3% |
Each distinct published measurement is retained. Same-snapshot canonical and configuration projections appear once.
| Published model / configuration | Score | Evidence & protocol | Source & dates |
|---|---|---|---|
Exact BenchLM registry variant Qwen3.6-27B; bulk export does not retain a complete upstream harness configuration.Qwen3.6 27BExact identitySource label without a registered configuration ID Canonical product: qwen3-6-27b | 97.8% | Source carriersource-checkedVersion & system2026 Source-native system | BenchLM public datasets — 2026-08-01Observed Checked |
Exact BenchLM registry variant LFM2.5-VL-450M; bulk export does not retain a complete upstream harness configuration.LFM2.5-VL-450MExact identitySource label without a registered configuration ID Canonical product: lfm2-5-vl-450m | 73.3% | Source carriersource-checkedVersion & system2026 Source-native system | BenchLM public datasets — 2026-08-01Observed Checked |
Exact BenchLM registry variant Qwen3.6-27B; bulk export does not retain a complete upstream harness configuration.Qwen3.6 27BExact identitySource label without a registered configuration ID Canonical product: qwen3-6-27b | 97.8% | Source carriersource-checkedVersion & system2026 Source-native system | BenchLM public datasets — 2026-07-27Observed Checked |
Exact BenchLM registry variant LFM2.5-VL-450M; bulk export does not retain a complete upstream harness configuration.LFM2.5-VL-450MExact identitySource label without a registered configuration ID Canonical product: lfm2-5-vl-450m | 73.3% | Source carriersource-checkedVersion & system2026 Source-native system | BenchLM public datasets — 2026-07-27Observed Checked |
Exact BenchLM registry variant Qwen3.6-27B; bulk export does not retain a complete upstream harness configuration.Qwen3.6 27BExact identitySource label without a registered configuration ID Canonical product: qwen3-6-27b | 97.8% | Source carriersource-checkedVersion & system2026 Source-native system | BenchLM public datasets — 21 July 2026Observed Checked |
Exact BenchLM registry variant LFM2.5-VL-450M; bulk export does not retain a complete upstream harness configuration.LFM2.5-VL-450MExact identitySource label without a registered configuration ID Canonical product: lfm2-5-vl-450m | 73.3% | Source carriersource-checkedVersion & system2026 Source-native system | BenchLM public datasets — 21 July 2026Observed Checked |
From result to context
A visual counting benchmark that tests whether a model can count objects and entities reliably in complex scenes.
The summary shows one best compatible source result per canonical product. Versions, effort settings and execution systems remain attached to the underlying records.
Scores from different versions or harnesses may not be interchangeable. The published source rows preserve those distinctions and their original units. A source result is not automatically an input to a current index.
This catalogue entry preserves available source evidence. Current indices admit only their specifically reviewed tracks and configurations.
Contamination risk: Unknown. Lifecycle: Active.