What the research found
Researchers stress-tested ChatGPT Health, OpenAI's consumer health tool launched in January 2026, using 60 clinician-authored vignettes across 21 clinical domains under 16 factorial conditions — 960 responses in total. Among gold-standard emergencies the system under-triaged 52% of cases, directing patients with diabetic ketoacidosis or impending respiratory failure to 24–48 hour evaluation rather than the emergency department, while handling classical emergencies such as stroke and anaphylaxis correctly. Failures followed an inverted U-shape, concentrated at the clinical extremes; when family or friends minimised symptoms, recommendations shifted toward less urgent care in edge cases (OR 11.7), and suicide-crisis safeguards activated unpredictably.
This is quiet harm in its purest form — no crash, no headline, just a patient told to wait 48 hours — and it happens in the consumer layer, outside every governance framework a provider controls.