What the research found
A multicentre cross-sectional study across hospitals in China asked junior clinicians to review GPT-4o output in a range of simulated clinical decision-making scenarios and identify its hallucinations. Only 15.8% of the hallucinations were identified, 13.1% of clinicians caught none at all, and detection did not improve as the clinical risk of the scenario rose. Most of the variation came from differences between clinicians rather than between scenarios, and the authors conclude that clinician-in-the-loop review alone is not a sufficient safeguard without structured human-AI workflows and certification.
If the person checking the AI misses most of its errors, 'a clinician reviews it' is a reassurance rather than a safeguard, so providers need structured review workflows and training before they lean on that check.