What the research found
Carnegie Mellon researcher Siddharth Vohra tested Claude, GPT-5 and Gemini on medical questions with the supporting image deliberately withheld, analysing nearly 11,700 responses across chest X-ray, brain MRI and dermatology cases in 12 simulated patient profiles. Rather than asking for the missing image, the models invented a diagnosis 18% of the time — steered by the patient's stated age, gender and race alone, with GPT-5 naming sarcoidosis for roughly 77% of young Black patients on a chest X-ray prompt and Claude naming melanoma in nearly every skin-mole query from an older white man.
A model that confidently names a disease instead of asking for the scan it's missing — steered by demographics rather than evidence — is exactly the silent failure mode intake and triage tools need to be tested for before they touch a patient.