PsyEval Benchmark Reveals AI Models Still Fall Short on Clinical Judgment and Empathy in Mental Health Care
核心洞察
A new comprehensive benchmark called PsyEval evaluated eleven advanced LLMs across knowledge, diagnosis, and emotional support dimensions in mental health contexts.
While models like Qwen2.5-72B (搜索) achieved 91.0% accuracy on Chinese medical exam questions, performance declined on urgent psychiatric scenarios and diagnostic tasks.
An "empathy gap" emerged: human counselors consistently outperformed AI in asking probing questions, with Exploration scores of 1.85–1.93 versus lower model scores.
A newly published benchmark study in npj Mental Health Research has put today's most advanced large language models through a rigorous mental health evaluation—and the results reveal that even the best AI systems still lack the emotional depth and clinical nuance required for reliable psychiatric care.
The benchmark, called PsyEval, assessed eleven LLMs—including GPT-4 (搜索), Qwen2.5-72B (搜索), and SoulChat (搜索)—across three core dimensions: knowledge understanding, diagnosis and assessment, and emotional support. While some models demonstrated strong factual recall and fluent language, the study uncovered an "empathy gap" compared to human counselors, significant prompt sensitivity, and a troubling safety-utility trade-off in diagnostic classification.
Knowledge proficiency varies by language and scenario
On factual knowledge tasks drawn from the US and Mainland China Medical Licensing Examinations (USMLE and MCMLE), larger general-purpose models performed well—but with notable language dependence. Qwen2.5-72B (搜索) achieved 91.0% accuracy on the Chinese MCMLE-mental dataset, while GPT-4 (搜索)-turbo led English tasks at approximately 76.0% accuracy.
However, performance declined when models faced urgent psychiatric scenario questions. On the USMLE crisis response dataset, GPT-4 (搜索)-turbo's accuracy fell to approximately 73.0%, suggesting that even strong factual knowledge does not guarantee reliable crisis decision-making.
The diagnostic paradox: safety versus accuracy
One of PsyEval's most striking findings was an "inverse scaling" phenomenon in diagnostic classification. Smaller models like LLaMa-3-8B achieved near-perfect accuracy classifying conditions such as ADHD (搜索) (100.0%) and anxiety (搜索) (96.0%), substantially outperforming GPT-4 (搜索), which achieved approximately 25% for ADHD.
The study authors hypothesized that these discrepancies reflect a safety-utility trade-off. Flagship foundation models may enforce guardrails that discourage them from diagnosing patients or practicing medicine, leading to refusals or ambiguous responses counted as wrong answers under PsyEval. Conversely, smaller, less guarded models may assign diagnostic labels more readily—potentially increasing the risk of overdiagnosis.
The empathy gap: humans still ask better questions
When comparing emotional support performance, the study revealed that while some leading LLMs matched or surpassed human counselors in language fluency and coherence, humans consistently outperformed AI in one critical area: asking probing questions to uncover deeper concerns.
Human counselors achieved an Exploration score of 1.85 on the Chinese PsyQA platform and 1.93 on the English Counsel-Chat dataset, consistently exceeding the evaluated AI models. This finding underscores that therapeutic communication involves more than fluent language—it requires genuine curiosity and emotional attunement that current models have not mastered.
Prompt engineering matters
Prompt style substantially influenced LLM performance. Scenario Simulation prompts—instructing models to adopt a mental health professional persona—generally produced the highest empathy scores. In contrast, Step-by-Step reasoning prompts often yielded more mechanical, less emotionally resonant responses.
Specialization trade-offs and study limitations
The findings from SoulChat (搜索), a domain-specific model, suggest that highly specialized conversational fine-tuning may come at the cost of general knowledge performance, though the authors caution against generalizing this observation to all specialized mental health models.
As a benchmark study, the research did not prospectively test models in patients, measure treatment outcomes, assess diagnostic safety in real-world settings, or evaluate autonomous crisis intervention. The authors emphasize that future model development should focus on incorporating culturally diverse alignment data and designing smarter safety mechanisms that protect users without unduly reducing clinical utility.
With depression (搜索) alone affecting 3.8% of the global population and treatment rates as low as 13.7% in lower-middle-income countries, the need for scalable mental health tools is urgent. PsyEval provides a critical framework for measuring whether AI can safely and effectively help meet that need—and the evidence so far suggests there is still considerable work ahead.
