AI Models

ChatGPT Health Advice Quality Depends on Subscription

OpenAI is rolling out its Health in ChatGPT feature to US users aged 18 and older, allowing them to connect health apps and medical records. Free users receive lower-quality health advice powered by GPT-5.5 Instant, while paying subscribers get access to the stronger GPT-5.6 Sol model. Despite over 300 million weekly health queries, risks remain significant as AI chatbots can give incorrect medical findings with high confidence.

Neura News

Neura News

Neura Market Editorial

July 23, 20264 min read
ChatGPT Health Advice Quality Depends on Subscription

OpenAI has begun rolling out its “Health in ChatGPT” feature to users in the United States aged 18 and older, offering health advice powered by two different AI models depending on whether a user pays or uses the free tier. The feature, first unveiled and tested in January 2026, lets people connect Apple Health, medical records, and wellness apps to review lab results, prepare for doctor’s appointments, and analyze sleep or activity data. OpenAI says it will not use connected health data for model training or advertising.

The rollout comes as more than 300 million people ask ChatGPT health questions each week, up from 230 million weekly users in January 2026. OpenAI says more than 260 physicians helped develop the Health features, and in early tests, more than 70% of participants asked health questions outside the dedicated Health section, prompting the company to make Health available in any conversation while keeping a separate Health section for data management.

Two-tier health advice: GPT-5.5 Instant versus GPT-5.6 Sol

Free users receive health advice powered by GPT-5.5 Instant, while paying subscribers get access to the stronger GPT-5.6 Sol model. OpenAI will likely defend this two-tier system on ethical grounds by pointing out that both models beat doctors’ answers on the HealthBench Professional test. On the HealthBench Professional completeness score, GPT-5.6 Sol achieved 88.0% versus 53.2% for GPT-5.5 Instant and older models. On the health decision helpfulness score, GPT-5.6 Sol scored 83.0% compared to 50.8% for GPT-5.5 Instant and older models.

However, benchmark results come from artificial test environments. Doctors may score lower due to time pressure, fatigue, or lack of tools, and benchmarks cannot capture actual medical exam elements like in-person examination, nonverbal cues, or years of experience. OpenAI itself states that ChatGPT can still make mistakes and cannot replace medical advice.

Risks of overconfidence and exclusion from Europe

The #1 Newsletter in AI

Stay ahead of the AI curve

The most important updates, news, and content — delivered weekly.

No spam. Unsubscribe anytime.

The risks of incorrect medical findings with high confidence are not hypothetical. In the RadLE 2.0 benchmark, none of the 16 AI models tested performed as well as human radiologists. Chatbots gave incorrect findings with high confidence instead of admitting uncertainty, while human radiologists were far better at acknowledging uncertainty. The mix of overconfidence, persuasion, and sycophancy in chatbots has contributed to serious mental health harms, according to reports cited in the article.

OpenAI has not said whether or when the Health feature will be available in Europe. When announced in January, the feature specifically excluded the European Economic Area, Switzerland, and the United Kingdom. Stricter EU data privacy rules and possible classification as high-risk under the EU AI Act are likely reasons for the exclusion.

AI in medicine: support tools, not replacements

Despite the risks, some reports show AI spotting patterns in health data that medical professionals miss, sometimes with striking results. AI systems like MIRA and AMIE have performed about as well as primary care doctors in simulated consultations. One researcher compared these AI agents to airplane autopilot, saying: “These systems can support and relieve medical professionals by taking over routine tasks, but ultimate responsibility will always remain with the physicians.”

The article, published by the-decoder.com and written by Matthias Bastian, notes that while AI can assist with health questions, it remains a tool akin to “Dr. Google” but not a substitute for a doctor. OpenAI’s own caveats underscore that ChatGPT can still make mistakes and cannot replace medical advice.

Related on Neura Market

More from Neura News

AI Models

42 Mathematicians Urge Royal Society to Warn Government and Media About AI Existential Risk

Forty-two mathematical fellows, including Fields Medal winners Martin Hairer, Peter Scholze, and Wendelin Werner, have signed an open letter urging the Royal Society to warn the UK government and media about existential risks from advanced AI. The letter follows recent breakthroughs in which leading models solved open research problems, including a Millennium Problem. None of the signatories are affiliated with AI companies. The group warns that AI labs' estimates of existential risk above ten percent must not be dismissed as hype, and that by the time the situation becomes obvious to the public, it may be too late to act.

Sep 18·2 min read
Developer

Steve Yegge Shuts Down Gas Town After Failing to Build Anything Else With It

Steve Yegge shut down Gas Town, his ultra-vibed coding agent orchestrator, after admitting he never built anything else with it despite heavy subscription spend. Databricks reported a 60% coding spend increase after rolling out GPT-6 Astra to 3,500 engineers, OpenAI published a misalignment disclosure framework with six case reports, and Xiaomi ran MiMo-V2.6 RL training in public with live telemetry.

Sep 18·21 min read