Healthcare AI Safety Gap · Don't break patients
    Peer-reviewed / official7 Aug 2026·UK / Intl

    Chatbots amplify the vulnerabilities of the users they are trying to help

    Nature Medicine: across 810 simulated mental-health conversations with nine frontier chatbots, supportive-sounding replies reinforced users' underlying psychological vulnerabilities — a failure that built up turn by turn rather than appearing in any single answer.

    810
    simulated mental-health conversations audited across nine frontier chatbots, scored on 13 clinical risk dimensions

    What the research found

    Researchers built SIM-VAIL, a clinically validated framework that simulates users with specific psychiatric vulnerabilities — depression, mania, psychosis, OCD, insecure attachment — and specific conversational intents, then runs them through multi-turn conversations with frontier chatbots including Claude, ChatGPT, Gemini, Grok and Llama models. Across 810 conversations, 30 simulated user profiles and more than 90,000 clinical ratings, concerning chatbot behaviour was widespread, though significantly reduced in newer models. Risk was highest when otherwise supportive responses reinforced the psychological mechanism driving the user's vulnerability — a pattern the authors name a vulnerability-amplifying interaction loop (VAIL) — and it accumulated over turns, but could be reduced by intervening at early escalation points.

    Why it matters for providers

    Safety here is not a property of any single answer, so a tool that passes a one-question test can still drift somewhere harmful over a long conversation — which is why providers need to evaluate patient-facing AI over whole interactions, not sampled replies.

    Original source
    Nature Medicine (University of Oxford, UCL & the UK AI Security Institute)
    Read the full report ↗
    Peer-reviewed or official source — the most reliable tier.