What the research found
Researchers built treatment-resistant-depression risk models from electronic health records at Mass General Brigham, Vanderbilt University Medical Center and Geisinger Clinic, covering patients first prescribed an antidepressant between 2004 and 2022. Internal validation was already weak (C-statistics 0.51–0.65); external validation at the other sites fell to 0.50–0.58, with precision-recall areas of 0.07–0.11 and low concordance between sites. The authors identify three failure mechanisms: heterogeneity in how treatment resistance is defined, biases baked into the source data, and site- and practice-level variation — and conclude that EHR data alone may not be enough for clinically meaningful risk stratification.
A model that works where it was built is not a model that works — and any health system buying a risk-stratification tool validated somewhere else should be asking for local evidence before it touches a patient.