Healthcare AI Safety Gap · Don't break patients
    Peer-reviewed / official20 Aug 2026·US

    Depression risk models barely beat a coin flip outside their home hospital

    npj Digital Medicine: treatment-resistant-depression risk models built on EHR data at three major US health systems failed to transport to each other, landing near chance.

    0.50–0.58
    external C-statistic — near-chance accuracy once the model moved to a new health system

    What the research found

    Researchers built treatment-resistant-depression risk models from electronic health records at Mass General Brigham, Vanderbilt University Medical Center and Geisinger Clinic, covering patients first prescribed an antidepressant between 2004 and 2022. Internal validation was already weak (C-statistics 0.51–0.65); external validation at the other sites fell to 0.50–0.58, with precision-recall areas of 0.07–0.11 and low concordance between sites. The authors identify three failure mechanisms: heterogeneity in how treatment resistance is defined, biases baked into the source data, and site- and practice-level variation — and conclude that EHR data alone may not be enough for clinically meaningful risk stratification.

    Why it matters for providers

    A model that works where it was built is not a model that works — and any health system buying a risk-stratification tool validated somewhere else should be asking for local evidence before it touches a patient.

    Original source
    npj Digital Medicine (Vanderbilt University Medical Center, Mass General Brigham, Geisinger)
    Read the full report ↗
    Peer-reviewed or official source — the most reliable tier.