Moonshot AI introduces PerceptionBench, a benchmark isolating visual perception in multimodal AI models. No tested frontier model reaches 60% accuracy, with GPT-5.6 Sol leading at 59.7%, followed by Kimi K3 at 58.5% and Claude Fable 5 at 57.2%. The benchmark's creators argue that many errors commonly attributed to logical reasoning failures actually occur during the initial image-reading stage, highlighting a persistent bottleneck in AI vision capabilities.
PerceptionBench visual perception benchmark
- ▪Moonshot AI published 3,000 tasks from an internal pool of over 17,000 verified questions for the PerceptionBench dataset, making the dataset and evaluation code available on GitHub.
- ▪The Chinese AI company Moonshot AI introduced PerceptionBench, a benchmark designed to isolate and test the visual perception of multimodal language models.
- ▪PerceptionBench evaluates visual perception across ten skill domains, including Visual Relation, Counting, Attributes, Depth & 3D, Localization, Comparison, Fine-grained Recognition, Context Integration, OCR, and Hallucination.
- ▪The authors of PerceptionBench built the benchmark's taxonomy by analyzing 42 open-source benchmarks to trace model errors back to the earliest failed step.
Multimodal AI model accuracy results
- ▪Open-source AI models trailed on the PerceptionBench benchmark, with Qwen3.5-397B-A17B scoring 47.5 percent and GLM-4.6V scoring 32.5 percent.
- ▪On the PerceptionBench evaluation, Kimi K3 scored 58.5 percent, Claude Fable 5 scored 57.2 percent, Gemini 3.1 Pro scored 56.2 percent, and GPT-5.5 scored 55.8 percent.
- ▪Hallucination, which tests whether models invent non-existent objects when the correct answer is zero, was the weakest skill on average across all tested models on PerceptionBench.
- ▪No tested frontier AI model achieved 60 percent accuracy on the PerceptionBench benchmark, with the highest score being 59.7 percent by GPT-5.6 Sol.
Perception failure versus reasoning errors
- ▪The creators of PerceptionBench argue that many multimodal AI model failures typically attributed to reasoning errors actually occur at the initial image-reading perception stage.
- ▪PerceptionBench isolates multi-step tasks into perception-only sub-questions to pinpoint which specific visual ability fails during model execution.
Related visual perception research
- ▪A study using the BabyVision benchmark showed frontier AI models fail at basic visual tasks tied to early childhood development, with Gemini 3 Pro scoring 49.7 percent compared to 94.1 percent for humans.
- ▪Prior research by the same Moonshot AI team on the WorldVQA benchmark showed the top model, Gemini 3 Pro, fell short of 50 percent accuracy at 47.4 percent.
Story comments
Loading comments…