SEP 11, 2026 · JOURNAL ARTICLE
Large language models for late-life depression: a blinded benchmark of clinical safety, geriatric appropriateness, and triage
Frontiers in Psychiatry
A rigorous blinded benchmarking study of three LLMs on a clinically relevant task (late-life depression triage) shows ChatGPT and Gemini perform moderately well on low-risk questions but all models fail substantially on high-risk scenarios, supporting cautious use only in low-risk educational contexts.
Reported
Clinically acceptable responses –…78.9%
Clinically acceptable responses –…72.2%
Clinically acceptable responses –…60.0%