AI Models

OpenAI's ChatGPT for Clinicians Beats Doctors on Benchmark

OpenAI released ChatGPT for Clinicians, a free tool for US healthcare workers. It scored 59.0 on the HealthBench Professional benchmark, higher than doctors' 43.7 despite their unlimited time and web access. The tool offers clinical searches, workflow templates, and CME credits, with strong safety ratings from tests.

Neura News

Neura News

Neura Market Editorial

April 23, 20263 min read
OpenAI's ChatGPT for Clinicians Beats Doctors on Benchmark

OpenAI's ChatGPT for Clinicians Beats Doctors on Benchmark

OpenAI introduced ChatGPT for Clinicians, a free chatbot tailored for medical professionals. This version targets verified physicians, nurses with advanced clinical qualifications, physician assistants, and pharmacists in the United States. The company also unveiled HealthBench Professional, a benchmark that tests AI on clinical tasks. OpenAI reports that GPT-5.4 in this setup scored 59.0 points, surpassing human doctors who scored 43.7 even with unlimited time and internet access.

Details of the HealthBench Professional Benchmark

HealthBench Professional evaluates AI in three areas: consultations, writing and documentation, plus medical research. It expands on the prior HealthBench and incorporates doctor-written conversations, multi-level physician scoring, and data filtering. OpenAI designed it to challenge models. Roughly one-third of examples stem from red teaming, where doctors sought model flaws. Difficult conversations appear 3.5 times more often.

GPT-5.4 within the ChatGPT for Clinicians environment achieved 59.0 overall. Human doctors managed 43.7 under ideal conditions. Other models trailed: base GPT-5.4 at 48.1, Anthropic's Claude Opus 4.7 at 47.0, Google's Gemini 3.1 Pro at 43.8, and xAI's Grok 4.2 at 36.1. The Clinicians version outperformed the standard GPT-5.4 by about 11 points. Questions remain on how much stems from the clinical environment versus benchmark design. Scores may not fully reflect real-world clinical use.

Development and Safety Evaluations

OpenAI created the benchmark and evaluated its models on it. The company cites independent checks like Stanford's MedHELM and MedMarks, where its models led. The benchmark and dataset sit publicly available. ChatGPT for Clinicians drew input from hundreds of medical advisors. Doctors tested 6,924 real clinical conversations before release. Karan Singhal from OpenAI's Health unit noted 99.6 percent of responses earned safe and accurate ratings.

In 355 cases with sources specified by three independent doctors, the tool referenced them more frequently than humans. Physicians reviewed over 700,000 model responses total. OpenAI positions the tool to assist clinicians, not supplant their decisions.

The #1 Newsletter in AI

Stay ahead of the AI curve

The most important updates, news, and content — delivered weekly.

No spam. Unsubscribe anytime.

Key Features for Clinical Use

Users get free access to OpenAI's latest frontier models. A clinical search draws from millions of peer-reviewed sources with real-time citations. A deep research option covers medical literature. Skills allow clinicians to build reusable templates for tasks like referral letters, prior authorizations, or patient instructions. Clinical research in the tool qualifies for continuing medical education credits in the US.

Privacy measures include no use of conversations for training. Optional HIPAA compliance via Business Associate Agreement supports protected health information.

Rollout and Broader Context

The launch limits to verified US clinicians. OpenAI plans global expansion and pilots with the Better Evidence Network abroad. It released a Health Blueprint for responsible AI in US healthcare.

AI use in medicine grows quickly. A 2026 American Medical Association survey showed 72 percent of US doctors using AI clinically, up from 48 percent in 2025. Millions of clinicians worldwide use ChatGPT weekly, with usage doubling yearly. Earlier in 2026, OpenAI offered ChatGPT for Healthcare to organizations for compliance and controls. Competitors like Anthropic, Microsoft, and Google advance medical AI, with Google emphasizing drug development via DeepMind.

OpenAI, founded in 2015, focuses on safe AGI development and leads in generative AI with ChatGPT since 2022. This clinician tool builds on that, targeting practical medical support.

Related on Neura Market

More from Neura News

Industry

Monday.com Joins Tech Layoff Trend Citing AI as Factor

Monday.com announced it will lay off about 20% of its workforce, or over 600 employees, citing a restructuring tied to its AI-driven growth strategy. The Tel Aviv-based work management software company joins a growing list of major tech firms, including Amazon, Meta, and Microsoft, that have cited artificial intelligence as a factor in job cuts this year. A new Financial Times analysis shows U.S. tech companies have slashed nearly 140,000 jobs since January, with AI often cited as a reason.

Jul 26·12 min read
General

Open-weight AI mirrors Kubernetes ecosystem shift

Tobi Knaup, co-founder of Mesosphere, draws parallels between the rise of Kubernetes and the current trajectory of open-weight AI models. He argues that open-weight models are becoming a neutral substrate for innovation, attracting a global ecosystem of developers, startups, and enterprises. The piece warns against US restrictions on Chinese open-weight models, advocating instead for American leadership through open releases, procurement strategies, and standards.

Jul 25·7 min read
General

Open-weight AI mirrors Kubernetes rise, US warned on bans

The author, a Mesosphere co-founder, draws parallels between the rise of Kubernetes and the current open-weight AI ecosystem. He argues that open-weight models are becoming a neutral platform for innovation, and warns that US restrictions on Chinese open-weight models could isolate American developers from a global ecosystem. The piece urges the US to compete by releasing frontier models, using procurement to create demand, building the stack, and setting standards rather than imposing bans.

Jul 25·7 min read
Industry

Power line failure reveals AI data center grid risks and solutions

A fallen power line near Washington, DC caused over 3 gigawatts of data center load to vanish from the PJM grid in seconds, spiking voltage across the region. The event, which made lights flicker from Northern Virginia to Chicago, highlights a growing problem as AI data centers become larger and more concentrated. Experts warn that without better coordination or technology like ON.Energy's battery-backed uninterruptible power supply, such disruptions will become more frequent and severe.

Jul 25·5 min read