Research

Harvard Study Shows AI Tops ER Doctors in Diagnoses

A Harvard-led study published in Science found OpenAI's o1 model delivered more accurate diagnoses than emergency room doctors for 76 real patient cases, especially during initial triage. Researchers from Harvard Medical School and Beth Israel Deaconess Medical Center compared AI outputs to physician assessments without preprocessing data. The results highlight AI's potential but stress the need for real-world trials and human oversight in critical decisions.

Neura News

Neura News

Neura Market Editorial

May 3, 20264 min read
Harvard Study Shows AI Tops ER Doctors in Diagnoses

Harvard Study Shows AI Tops ER Doctors in Diagnoses

Researchers from Harvard Medical School and Beth Israel Deaconess Medical Center have released a study that tests large language models in medical scenarios. The work, published this week in the journal Science, includes analysis of actual emergency room cases. In those instances, one AI model provided diagnoses that matched or exceeded the accuracy of human doctors.

The team, made up of physicians and computer scientists, ran multiple experiments. They compared OpenAI's o1 and 4o models against human physicians. The goal was to gauge performance across different medical tasks.

Experiment with Real ER Patients

A key part of the research involved 76 patients who visited the Beth Israel emergency room. Two attending physicians reviewed each case and offered diagnoses. Separately, OpenAI's o1 and 4o models generated their own diagnoses based on the same records.

Two additional attending physicians then evaluated all these diagnoses. They did not know which came from AI and which from humans. This blind setup ensured fair comparison.

The study reported that at every stage of diagnosis, o1 either matched or slightly outperformed the two original physicians. The 4o model also performed well. Differences stood out most at the initial triage stage. There, limited patient information creates high pressure for quick, correct calls.

Strong Triage Performance

During triage, the o1 model achieved an exact or very close diagnosis in 67% of cases. One physician reached that level in 55% of instances. The other physician succeeded in 50%.

Importantly, the AI received no special preparation. Models saw only the text from electronic medical records available at each decision point. This matched exactly what physicians used.

Arjun Manrai, who leads an AI lab at Harvard Medical School and served as a lead author, commented in a Harvard press release. "We tested the AI model against virtually every benchmark, and it eclipsed both prior models and our physician baselines," he said.

Harvard Medical School traces its roots to 1782. It stands as one of the world's leading institutions for medical education and research. Beth Israel Deaconess Medical Center, a major teaching hospital in Boston, affiliates closely with Harvard. It handles thousands of emergency visits yearly and supports advanced clinical studies.

The #1 Newsletter in AI

Stay ahead of the AI curve

The most important updates, news, and content — delivered weekly.

No spam. Unsubscribe anytime.

OpenAI, the developer of o1 and 4o, focuses on artificial general intelligence. The company released these models as advanced tools for complex reasoning tasks. O1 emphasizes step-by-step thinking, while 4o handles multimodal inputs including text and images.

Cautions and Next Steps

The study makes no claim that AI can handle life-or-death choices in emergency rooms yet. Instead, results point to an urgent need for prospective trials. Those would test the technology directly in patient care environments.

Researchers limited their work to text-based inputs. They noted that foundation models struggle more with images or other non-text data, based on prior studies.

Adam Rodman, a Beth Israel physician and lead author, spoke to the Guardian. He pointed out the lack of a formal framework for AI accountability in diagnoses. Patients, he said, prefer human guidance for life-or-death and tough treatment choices.

This research builds on growing interest in AI for healthcare. Institutions like Harvard continue to explore how such tools might assist clinicians without replacing them.

Broader Implications for Medical AI

The findings come amid rapid advances in AI capabilities. Models like those from OpenAI now tackle specialized domains such as medicine. Yet, deployment requires rigorous validation.

Beth Israel Deaconess, founded in 1916 through mergers, ranks among top U.S. hospitals for research. It pioneered electronic health records and now integrates AI in trials.

Manrai's AI lab at Harvard examines computational methods for health data analysis. Rodman's clinical role complements this with practical emergency medicine insights.

Overall, the study underscores AI's promise in high-stakes diagnostics. It also reinforces the value of human expertise and the push for controlled real-world testing.

Related on Neura Market

More from Neura News

AI Models

Google Unveils Gemini 3.6 Flash, 3.5 Flash-Lite, and Cyber Model

Google has released three new Gemini models: 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber. The 3.6 Flash model offers improved coding and knowledge work with 17% fewer output tokens and lower costs. The 3.5 Flash-Lite is the fastest in the series at 350 tokens per second, designed for high-throughput agentic tasks. The 3.5 Flash Cyber model, available only to governments and trusted partners via CodeMender, focuses on finding and fixing cybersecurity vulnerabilities. Google also noted that Gemini 3.5 Pro is being tested with partners and that pre-training for Gemini 4 has begun.

Jul 21·5 min read
AI Models

Alibaba Qwen-Image-3.0 renders infographics and tiny text in one pass

Alibaba's Qwen team released Qwen-Image-3.0, an image generator designed for practical applications like newspaper layouts and complex infographics. The model processes prompts of up to 4,500 tokens and can render legible text as small as ten pixels, mathematical formulas, and twelve languages in a single pass. It is currently available through invite-only API access, with plans to integrate it into first-party apps like Qwen Chat soon.

Jul 21·4 min read
AI Models

Google Unveils Gemini 3.6 Flash, 3.5 Flash-Lite, and Cyber Model

Google DeepMind has introduced three new Gemini models: 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber. The 3.6 Flash model offers improved coding and multimodal performance with 17% fewer output tokens and lower cost. The 3.5 Flash-Lite is the fastest in its series at 350 output tokens per second, designed for high-throughput agentic tasks. The 3.5 Flash Cyber, fine-tuned for cybersecurity, will be available exclusively to governments and trusted partners via the CodeMender agent.

Jul 21·6 min read