A new study utilizing a "Shadow Evaluation" methodology reveals that frontier AI agents, including Claude Opus 4.8 and GPT-5.6 Sol, fail to independently produce publishable scientific research. While the agents successfully managed complex engineering tasks like debugging GPU code and compiling LaTeX papers, human reviewers rejected all agent-generated papers due to poor scientific judgment, bizarre experimental choices, and unreadable prose. These findings directly challenge autonomous research claims made by major labs like Anthropic and OpenAI.
Oct 1, 2026 · 4 sources
Oct 1, 2026 · 2 sources
Sep 29, 2026 · 14 sources
Sep 29, 2026 · 7 sources
Story comments
Loading comments…