Healthcare AI Safety Gap · Don't break patients
    Peer-reviewed / official24 Sep 2026·China

    Junior doctors caught fewer than one in six AI hallucinations

    npj Digital Medicine: in a multicentre study across Chinese hospitals, junior clinicians reviewing GPT-4o's output in simulated clinical scenarios identified only 15.8% of its hallucinations.

    15.8%
    of GPT-4o hallucinations identified by junior clinicians in simulated clinical decision-making

    What the research found

    A multicentre cross-sectional study across hospitals in China asked junior clinicians to review GPT-4o output in a range of simulated clinical decision-making scenarios and identify its hallucinations. Only 15.8% of the hallucinations were identified, 13.1% of clinicians caught none at all, and detection did not improve as the clinical risk of the scenario rose. Most of the variation came from differences between clinicians rather than between scenarios, and the authors conclude that clinician-in-the-loop review alone is not a sufficient safeguard without structured human-AI workflows and certification.

    Why it matters for providers

    If the person checking the AI misses most of its errors, 'a clinician reviews it' is a reassurance rather than a safeguard, so providers need structured review workflows and training before they lean on that check.

    Original source
    npj Digital Medicine (multicentre study; affiliated hospitals in Wuxi, Harbin, Jiamusi and Qiqihar, China)
    Read the full report ↗
    Peer-reviewed or official source — the most reliable tier.