Geo News
Community curated by people like you
LatestAICryptoHealthWorld AffairsUS Politics
New benchmark shows AI models struggle with visual perception, none reach 60% accuracy
00

New benchmark shows AI models struggle with visual perception, none reach 60% accuracy

Aug 15, 2026

Moonshot AI introduces PerceptionBench, a benchmark isolating visual perception in multimodal AI models. No tested frontier model reaches 60% accuracy, with GPT-5.6 Sol leading at 59.7%, followed by Kimi K3 at 58.5% and Claude Fable 5 at 57.2%. The benchmark's creators argue that many errors commonly attributed to logical reasoning failures actually occur during the initial image-reading stage, highlighting a persistent bottleneck in AI vision capabilities.

PerceptionBench visual perception benchmark

  • ▪Moonshot AI published 3,000 tasks from an internal pool of over 17,000 verified questions for the PerceptionBench dataset, making the dataset and evaluation code available on GitHub.
  • ▪The Chinese AI company Moonshot AI introduced PerceptionBench, a benchmark designed to isolate and test the visual perception of multimodal language models.
  • ▪PerceptionBench evaluates visual perception across ten skill domains, including Visual Relation, Counting, Attributes, Depth & 3D, Localization, Comparison, Fine-grained Recognition, Context Integration, OCR, and Hallucination.
  • ▪The authors of PerceptionBench built the benchmark's taxonomy by analyzing 42 open-source benchmarks to trace model errors back to the earliest failed step.

Multimodal AI model accuracy results

  • ▪Open-source AI models trailed on the PerceptionBench benchmark, with Qwen3.5-397B-A17B scoring 47.5 percent and GLM-4.6V scoring 32.5 percent.
  • ▪On the PerceptionBench evaluation, Kimi K3 scored 58.5 percent, Claude Fable 5 scored 57.2 percent, Gemini 3.1 Pro scored 56.2 percent, and GPT-5.5 scored 55.8 percent.
  • ▪Hallucination, which tests whether models invent non-existent objects when the correct answer is zero, was the weakest skill on average across all tested models on PerceptionBench.
  • ▪No tested frontier AI model achieved 60 percent accuracy on the PerceptionBench benchmark, with the highest score being 59.7 percent by GPT-5.6 Sol.

Perception failure versus reasoning errors

  • ▪The creators of PerceptionBench argue that many multimodal AI model failures typically attributed to reasoning errors actually occur at the initial image-reading perception stage.
  • ▪PerceptionBench isolates multi-step tasks into perception-only sub-questions to pinpoint which specific visual ability fails during model execution.

Related visual perception research

  • ▪A study using the BabyVision benchmark showed frontier AI models fail at basic visual tasks tied to early childhood development, with Gemini 3 Pro scoring 49.7 percent compared to 94.1 percent for humans.
  • ▪Prior research by the same Moonshot AI team on the WorldVQA benchmark showed the top model, Gemini 3 Pro, fell short of 50 percent accuracy at 47.4 percent.

1 source

The Decoder
New benchmark confirms AI models still perform poorly at visual perception
View source article

Featured stories

View more in AI research & benchmarks

AMD acquires AI startup World Labs founded by Fei-Fei Li for $8.2 billion

Sep 28, 2026 · 7 sources

Sarvam AI releases Saaras V4 speech model covering 22 Indian languages

Sep 26, 2026 · 2 sources

Story comments

Loading comments…

Related Projects

Moonshot AI

Topics

AI research & benchmarksAI safety benchmarksDomain specific AI benchmarksMultimodal modelsModel evaluation methodology

Featured stories

View more in AI research & benchmarks

AMD acquires AI startup World Labs founded by Fei-Fei Li for $8.2 billion

Sep 28, 2026 · 7 sources

Sarvam AI releases Saaras V4 speech model covering 22 Indian languages

Sep 26, 2026 · 2 sources