What the research found
Reviewing the evidence on how general-purpose chatbots and AI companions handle suicide risk disclosures, researchers at Crisis Text Line found that nearly all existing safety tests rely on single-turn, researcher-scripted prompts rather than the indirect, gradually-disclosed way teens actually talk — and that chatbot performance degrades sharply on exactly those realistic patterns. No study has yet measured what happens to real teens' suicide risk after they talk to a chatbot. The authors lay out a research agenda to close that gap.
The riskiest population is using the least-tested product feature, and providers still have no real-world evidence of harm or benefit to act on — only lab-style tests that don't resemble how a struggling teen actually opens up.