AI Outperforms Emergency Room Physicians in Diagnostic Accuracy in Harvard Study
A Harvard Medical School and Beth Israel Deaconess Medical Center study published May 6, 2026 found OpenAI's o1-preview surpassed two internal medicine attending physicians in emergency room diagnostic accuracy across virtually every benchmark. Two attending physicians blindly assessed competing diagnoses without knowing which came from AI or humans, with results favoring AI. A December 2025 study found 67% of physicians changed treatment decisions after AI suggested the opposite.
Harvard diagnostic accuracy study
▪Adam Rodman is a senior author of the Harvard Medical School and Beth Israel Deaconess Medical Center AI diagnostic study and a Beth Israel doctor.
▪The article about AI outperforming emergency room physicians in diagnostic accuracy was published on May 6, 2026.
▪OpenAI's o1-preview AI model eclipsed both prior AI models and physician baselines against virtually every benchmark in the Harvard Medical School study.
▪The Harvard Medical School and Beth Israel Deaconess Medical Center study results favored AI over human physicians in emergency room diagnoses.
▪Arjun Manrai is a senior co-author of the Harvard Medical School and Beth Israel Deaconess Medical Center AI diagnostic study and an assistant professor of biomedical informatics at Harvard's Blavatnik Institute.
▪A Harvard Medical School and Beth Israel Deaconess Medical Center study compared emergency room diagnoses from OpenAI's o1-preview against diagnoses offered by two internal medicine attending physicians.
▪Peter Brodeur is a co-author of the Harvard Medical School and Beth Israel Deaconess Medical Center AI diagnostic study and a Harvard clinical fellow in medicine at Beth Israel Deaconess.
▪Two attending physicians assessed the competing diagnoses from OpenAI's o1-preview and human physicians without knowing which results were from humans or AI.
AI adoption in healthcare
▪Google DeepMind's Alphafold is advancing biological research.
▪An AI system in Utah is prescribing medicine to patients without a physician in the loop.
▪Physicians have noted that the Utah AI system prescribing medicine without a physician in the loop could put patients at risk.
▪Some emergency rooms have deployed generative AI to take notes and create medical records.
▪Fortune Brainstorm Tech will take place in Aspen from June 8-10, 2026, marking 25 years of the conference.
Study methodology with raw data
▪The Harvard Medical School and Beth Israel Deaconess Medical Center study presented each case exactly as it appeared in an electronic health record without cleaning up the data.
▪AI models are consistently scoring close to 100% on multiple-choice tests used to evaluate medical AI models.
Physician decision-making impact
▪There is no formal framework for accountability when it comes to AI diagnoses.
▪A December 2025 study found that 67% of physicians who initially recommended against treatment for a patient changed their decision after AI suggested the opposite.
▪AI is good at diagnosing but tends to suggest unnecessary testing that could do more harm than good.
Patient trust concerns
▪Patients want humans to guide them through life or death decisions and challenging treatment decisions.
Perspective of Medical experts and physicians
▪AI diagnostic technology tends to suggest unnecessary testing that could do more harm than good despite strong diagnostic performance.
Story comments
Loading comments…