Researchers led by internist and clinical AI researcher Adam Rodman published a compilation of experiments in Science on Thursday demonstrating that an OpenAI large language model outperformed physicians on case-based diagnostic and clinical reasoning evaluations, including one experiment using real-world data from a Boston emergency department. The study represents one of the largest AI-physician comparison studies to date, building on methodology from a 1959 Science paper that established criteria for evaluating clinical decision support systems. However, co-senior author Rodman warns the results should not be misconstrued as proof of AI safety and efficacy for real patient treatment, noting all experiments used simulated and historical cases rather than real-time scenarios, even as generative AI tools are being heavily marketed to patients and clinicians.
Aug 10, 2026 · 3 sources
Aug 8, 2026 · 2 sources
Aug 7, 2026 · 3 sources
Aug 7, 2026 · 2 sources
Story comments
Loading comments…