Machine Learning in Healthcare · Journal article
Translational Psychiatry · July 10, 2026
Early or partial results. Treat as a signal, not a conclusion.
This is a retrospective validation study of eight deep-learning models applied to clinical notes from a large healthcare EHR to classify antidepressant treatment response. The best-performing model achieved an AUROC of 0.88 and PPV of 0.84, but the phenotype depends on documented clinical observations rather than standardised depression assessments, and no prospective clinical utility or external validation is reported.
Retrospective cohort study with manual annotation and machine-learning model validation. Adults with depression and a co-occurring antidepressant prescription from Mass General Brigham healthcare system EHR spanning 1990–2018; n=111,572 total, with 4,299 manually reviewed and annotated.. Intervention: Eight deep-learning-based natural language processing models trained to classify antidepressant treatment response from clinical notes.. Compared with: Manual review and classification of clinical notes as gold standard for phenotyping; no comparison between deep-learning and alternative automated or clinical approaches reported.. n = 111,572. Mass General Brigham healthcare system (United States).
All eight deep-learning models achieved areas under the receiver operator curve (AUROC) of at least 0.80 Best-performing model (Longformer-large with sliding window) achieved AUROC = 0.88 and PPV = 0.84 at specificity of 0.88 Positive predictive values across models ranged from 0.72 to 0.91
Safety was not reported in the material analysed. Check the source before drawing any conclusion about harm.
Clinicians and researchers should recognize this as a proof-of-concept for automated classification of documented treatment response from EHR notes, not as validated clinical decision support. Prospective validation against standardised outcome measures and external validation in other health systems would be needed before clinical deployment.
A single-centre, real-world validation study of machine-learning phenotyping using surrogate (note-based) classification of treatment response, without comparison to gold-standard clinical outcomes or prospective clinical utility demonstration.
As stated by the source record.
Quoted from the source exactly as published.
Clinicians and researchers should recognize this as a proof-of-concept for automated classification of documented treatment response from EHR notes, not as validated clinical decision support. Prospective validation against standardised outcome measures and external validation in other health systems would be needed before clinical deployment.
Graded across the dimensions that decide whether you should act, each from what the source actually supports. There is no single score, and where a dimension was not assessed it says so.
ABSTRACT Efficient, accurate phenotyping for antidepressant treatment response in electronic health records (EHRs) could facilitate precision psychiatry applications but remains a challenge. Increasingly, artificial intelligence methods using “deep learning” applied to clinical data have shown promise in complex classification problems. Here, we systematically evaluate the performance of eight deep-learning-based natural language processing models in classifying response to antidepressants in a large real-world healthcare setting. We obtained data spanning 1990-2018 for adults with depression and a co-occurring antidepressant prescription from the EHR data warehouse of the Mass General Brigham healthcare system (n=111,572). Clinical notes were collected for the following time windows after antidepressant initiation: (1) 2 days to 4 weeks, (2) 4–12 weeks, and (3) 12–26 weeks. A stratified random sample of these note sets (total 4,299 across time periods) were manually reviewed to classify response status as “improved” or “no evidence of improvement” in depression symptoms. All models performed well, with areas under the receiver operator curve (AUROC) of at least 0.80. Positive predictive values (PPVs) ranged from 0.72 – 0.91. In general, models incorporating more information-dense and longer text sequences performed better than others. The best performing model (Longformer-large with sliding window) had an AUROC = 0.88 and PPV = 0.84 at a specificity of 0.88. Our results indicate that deep learning methods applied to EHR data can accurately classify antidepressant response in a real-world healthcare setting. Automated treatment response classification may facilitate a range of research and clinical decision support applications.
Taken from the source record, never inferred. Follow any of these and new work involving them reaches your briefing.