Geo News
Community curated by people like you
LatestAICryptoHealthWorld AffairsUS Politics

Model evaluation methodology stories

Aug 15, 2026

New benchmark shows AI models struggle with visual perception, none reach 60% accuracy

Moonshot AI's PerceptionBench reveals that leading multimodal AI models, including GPT-5.6 Sol, perform poorly at basic visual perception tasks when separated from logical reasoning, with no frontier model achieving 60% accuracy. The benchmark demonstrates that many errors attributed to reasoning actually occur during the image-reading stage.

Aug 15, 2026·1 source
00
Jul 24, 2026

Snorkel AI Highlights First Wave of Open Benchmarks Grants Projects

Snorkel AI has revealed the initial projects receiving support from its Open Benchmarks Grants, a $3 million initiative dedicated to funding open-source datasets, benchmarks, and evaluation tools.

Jul 24, 2026·1 source
00

Top claims

  • ▪Prior research by the same Moonshot AI team on the WorldVQA benchmark showed the top model, Gemini 3 Pro, fell short of 50 percent accuracy at 47.4 percent.
  • ▪Moonshot AI published 3,000 tasks from an internal pool of over 17,000 verified questions for the PerceptionBench dataset, making the dataset and evaluation code available on GitHub.
  • ▪On the PerceptionBench evaluation, Kimi K3 scored 58.5 percent, Claude Fable 5 scored 57.2 percent, Gemini 3.1 Pro scored 56.2 percent, and GPT-5.5 scored 55.8 percent.

Subtopics

AI research & benchmarks2AI safety benchmarks1AI startups1Domain specific AI benchmarks1Multimodal models1Open-source AI1

Related timelines

AI Data Center Gold Rush

101 stories

Congress

108 stories

Crypto hacks

100 stories

Ebola outbreak

58 stories

Iran War

209 stories