A person with asthma types their symptoms into ChatGPT Health. The model reads them correctly. In its own written explanation it identifies the early warning signs of respiratory failure, names them, and understands what they mean. Then it tells the person to wait and see how things go.
That single interaction, recorded in the first independent safety evaluation of the tool, is the most alarming thing in the study, and not for the reason it first appears. The model was not confused. It got the medicine right and the instruction wrong. Whatever failed there was not knowledge.
What The Study Actually Did
The evaluation was published in Nature Medicine on February 23, 2026, and it is the first independent safety assessment of ChatGPT Health rather than a vendor self-report. That distinction carries most of the weight here. The lead author is Dr. Ashwin Ramaswamy, an instructor of urology at the Icahn School of Medicine at Mount Sinai.
Researchers did not run a handful of gotcha prompts. They built 60 structured clinical scenarios spanning 21 medical specialties, then ran each one under 16 different contextual conditions, varying how the patient presented themselves and their situation. That produces 960 separate interactions with the system. This is a real methodology, and it is the kind of thing that has been conspicuously missing from most public argument about medical AI.
The Triage Number
Across cases where physicians independently determined the patient needed emergency care, ChatGPT Health under-triaged more than half of them. It told people who needed an emergency department to stay home or book a routine appointment instead.
Sit with the base rate for a moment. If you flipped a coin to decide whether to send someone to the hospital, you would perform about as well on this specific measure. A tool that produces fluent, confident, well-organized medical prose is delivering emergency routing at roughly the accuracy of chance.
Where It Worked, Which Matters More Than It Sounds
Now the part that complicates the outrage, and it deserves to be stated plainly rather than buried. ChatGPT Health performed well on textbook emergencies. Stroke presentations, severe allergic reactions, the cases that appear in the first chapter of every clinical reference with unmistakable red-flag symptoms. On those, it did its job.
What it struggled with were nuanced situations where the danger is not immediately obvious. That is a coherent and even predictable failure profile for a system trained on text: it recognizes the canonical description of a stroke because the canonical description of a stroke appears in its training data ten thousand times. It misses the atypical presentation because atypical presentations are, by definition, the ones that do not match the pattern.
All of which is precisely backwards from what a triage tool is for. Nobody needs an AI to tell them that sudden facial droop and slurred speech means call an ambulance. The entire value of a triage assistant lives in the ambiguous cases, the ones where a reasonable person genuinely cannot tell whether this is serious. That is the exact region where the tool degrades.
The Inversion
The suicide safety findings are the ones that should end the conversation about deploying this at consumer scale in its current form.
ChatGPT Health was built to surface the 988 Suicide and Crisis Lifeline when a user appears to be at high risk. Investigators found that these alerts appeared inconsistently. That alone would be a serious defect. What they actually found is worse than inconsistency. Dr. Nadkarni described the pattern as alerts that were inverted relative to clinical risk, appearing more reliably for lower-risk scenarios than for cases where someone shared how they intended to hurt themselves.
Read that again with the ordering in mind. Vague distress got the hotline. A specific plan for self-harm frequently did not.
An unreliable safety net is a bad product. An anti-correlated safety net is a different category of thing, because it actively teaches the user the wrong lesson. Someone who receives a crisis resource during a low-risk conversation learns that the system watches for danger and speaks up. That learned trust is then carried into the conversation where they disclose a plan and the system says nothing. The silence reads as reassurance. The tool's earlier correct-seeming behavior is what makes its later failure dangerous.
Why "False Sense Of Security" Is The Precise Complaint
Alex Ruani of University College London called the findings unbelievably dangerous and pointed at the mechanism: the tool creates a false sense of security. That phrase is doing more work than it looks like.
Consider the counterfactual honestly. A person with worrying symptoms and no ChatGPT Health has an uncomfortable, unresolved feeling. That discomfort is useful. It is what eventually drives someone to call a nurse line, text a friend who works in healthcare, or drive to urgent care at eleven at night. The same person who consults a fluent AI and is told to monitor symptoms and book an appointment next week does not have that discomfort anymore. It has been resolved, articulately, by something that sounded like it knew.
Harm here is not merely that the tool gave bad advice. It is that it removed the anxiety that would otherwise have produced good behavior. A confident wrong answer is worse than no answer, and this is the cleanest documented example of that principle we have seen.
What Would Actually Fix This
The engineering response here is not especially exotic, which is part of what makes the current state frustrating.
Asymmetric costs demand asymmetric thresholds. Under-triage and over-triage are not equivalent errors. Sending someone to an emergency department unnecessarily costs money and hours. Not sending someone who needed to go can cost a life. A system that treats those as symmetric classification errors is optimizing the wrong objective, and the fix is to bias hard toward escalation in every ambiguous case.
Crisis-line triggering should not be a judgment call at all. Any mention of self-harm intent, method, or plan should surface 988 deterministically, outside the model, as a hard-coded rule that no generated text can suppress. The moment that decision lives inside a probabilistic system, it inherits that system's variance, and this study is what that variance looks like.
And the model's own reasoning should be usable as a signal. In the asthma case, the text contained the phrase that should have triggered escalation. The system wrote the words "respiratory failure" and then recommended waiting. A rule that escalates whenever the generated explanation names a critical condition would have caught that specific failure without any change to the underlying model.
The Part Nobody Wants To Say
There is a version of this article that concludes medical AI is a fraud and everyone should stop. That is not what the evidence supports, and pretending otherwise would be its own kind of dishonesty. The tool handled canonical emergencies competently. Millions of people have no realistic access to a clinician at two in the morning, and for them the alternative to imperfect AI triage is often nothing at all.
But that argument only works if the system fails safe, and this one fails dangerous. It under-triages the ambiguous cases and goes quiet on the specific plans. Both failures point the same direction: away from help, at exactly the moments help matters most.
Sixty scenarios. Twenty-one specialties. Nine hundred sixty interactions. One independent study, and the safety feature was running backwards. File it next to the running timeline of AI failures and the docket of AI litigation, because the next product to ship with an inverted safety alert is being demoed to somebody this week.