What the research found
Researchers built SIM-VAIL, a clinically validated framework that simulates users with specific psychiatric vulnerabilities — depression, mania, psychosis, OCD, insecure attachment — and specific conversational intents, then runs them through multi-turn conversations with frontier chatbots including Claude, ChatGPT, Gemini, Grok and Llama models. Across 810 conversations, 30 simulated user profiles and more than 90,000 clinical ratings, concerning chatbot behaviour was widespread, though significantly reduced in newer models. Risk was highest when otherwise supportive responses reinforced the psychological mechanism driving the user's vulnerability — a pattern the authors name a vulnerability-amplifying interaction loop (VAIL) — and it accumulated over turns, but could be reduced by intervening at early escalation points.
Safety here is not a property of any single answer, so a tool that passes a one-question test can still drift somewhere harmful over a long conversation — which is why providers need to evaluate patient-facing AI over whole interactions, not sampled replies.