Geo News
Community curated by people like you
LatestAICryptoHealthWorld AffairsUS Politics

Open model benchmarks & leaderboards stories

Aug 13, 2026

Ling 3.0 Flash becomes top-performing open model in its size class

Ling 3.0 Flash achieved a score of 38 points on the Artificial Analysis Intelligence Index, matching Qwen3.6 performance and representing a significant improvement over its predecessor in the open-source AI model category.

Aug 13, 2026·1 source
00
Apr 14, 2026

MiniMax Open Sources M2.7 Self-Evolving Agent Model Achieving 56.22% on SWE-Pro Benchmark

MiniMax released model weights for MiniMax M2.7 on Hugging Face, a self-evolving agent model that scores 56.22% on SWE-Pro and 57.0% on Terminal Bench 2. The company also released MMX-CLI, a command-line interface providing native access to image, video, speech, music, vision, and search capabilities.

Apr 14, 2026·2 sources
00
Apr 12, 2026

Arcee AI Spends Half Its Venture Capital to Build Open Reasoning Model Rivaling Claude Opus

US startup Arcee AI invested approximately half of its total venture capital to train Trinity-Large-Thinking, a 400 billion parameter open reasoning model designed to compete with Anthropic's Claude Opus in agent tasks.

Apr 12, 2026·1 source
00
Apr 8, 2026

Z.AI Releases GLM-5.1 Open-Source Model Capable of 8-Hour Autonomous Execution

Chinese AI company Z.AI unveiled GLM-5.1, a 754B parameter open-source agentic model that achieves state-of-the-art performance on SWE-Bench Pro and can run autonomously for up to 8 hours on long-horizon engineering tasks.

Apr 8, 2026·2 sources
00

Top claims

  • ▪On per-token pricing, Ling 3.0 Flash is cheaper than any comparably capable model as of August 13, 2026.
  • ▪Ling 3.0 Flash refuses to answer questions for which it lacks reliable answers more frequently than its predecessor model.
  • ▪Ling 3.0 Flash remains cheaper on a per-task basis than Qwen3.6 27B despite burning through more tokens on complex tasks.

Subtopics

Open-source AI4AI agents3Open model families3China AI regulations2Large language models (LLMs)2Software engineering benchmark (SWE-bench)2AI research & benchmarks1AI startups1AI tools & products1Multimodal models1Reasoning models1

Related timelines

Payments

118 stories

Trump administration

116 stories

Congress

108 stories

Iran War

209 stories

Russia-Ukraine war

136 stories