JUL 7, 2026 · JOURNAL ARTICLE
Benchmarking large language models against practicing clinicians on psychopathological assessment
npj Digital Medicine
Proof-of-concept study using simulated interviews and expert consensus reference rather than clinical outcomes; demonstrates feasibility but requires validation in real patient settings before clinical application.
Reported
Top LLM accuracy (GPT-5.1, Gemini…0.72
GPT-5.1 accuracy percentile relat…64th percentile
GPT-5.1 accuracy – depression sce…0.81