AI Chatbot Risk in Mental Health: New Simulation Framework Quantifies Harm Potential Across Models
核心洞察
A new study in BMJ Mental Health uses AI-generated simulations to estimate the risk of harm when patients with mental health conditions use AI chatbots, finding risk varies by roughly four orders of magnitude depending on the model and patient context.
Researchers tested 14 open-source AI models on three safety tasks—recognizing suicidal ideation (搜索), therapy requests, and therapy-like conversation drift—with two psychiatrists validating all synthetic exchanges.
The tools were weakest at detecting subtle suicidal ideation (搜索), including ambivalent or conditional language, mirroring the areas that demand the most clinical skill.
A patient tells her therapist she has been using a general-purpose AI chatbot late at night when she feels anxious. The clinician already knows the tragic headlines about chatbots saying things they never should have. But sitting across from this particular patient, the question is narrower and far harder: How likely is it that this tool, used by this person, in this situation, will lead to harm?
A new study from the Division of Digital Psychiatry at Beth Israel Deaconess Medical Center (BIDMC), published in BMJ Mental Health, offers a step toward a real answer—by using the AI chatbots themselves to help estimate the risk for a specific patient with a specific illness.
Why "Is AI Safe?" Is the Wrong Question
Medical devices such as pacemakers have been regulated for decades using a straightforward chain of logic: a hazard is anything that could cause harm, a hazardous situation is when a patient is actually exposed to that hazard, and harm is the bad outcome that results. Regulators like the FDA ask manufacturers to estimate the steps that turn hazards into harm, the likelihood of harm, its severity, and whether the benefits outweigh the risks.
That framework was designed for devices that behave predictably. A large language model (LLM), by definition, does not. LLMs are probabilistic; their behavior shifts with each update. Complicating matters further, the exact same output can be harmless to one person and dangerous to another—or even the same person at different times. This is why it is hard to say they are always harmful or always safe.
Therapists are already familiar with the cost of black-and-white thinking. Instead of "always harmful" or "always safe," what they need is a realistic estimate—one that does not shy away from admitting its uncertainty.
Turning Failures Into Numbers
The study's core idea is to use simulation to fill the gaps where real-world data are absent. Because researchers cannot expose real patients to thousands of risky conversations to see what happens, the team had chatbots generate the conversations instead. One AI system produced large numbers of synthetic user statements and chatbot exchanges. Then, two psychiatrists reviewed every single one to confirm it actually represented what it claimed to, such as suicidal ideation (搜索) or a request for therapy.
The researchers then tested 14 different open-source AI models on three safety-relevant tasks: recognizing suicidal ideation (搜索), recognizing when a user is asking for therapy, and recognizing when a conversation has drifted into therapy-like territory. How often each model missed these things yielded two estimates: the probability that a hazard becomes a hazardous situation, and the probability that the situation goes on to cause harm.
The headline finding is not a single safety score—it is how wide the range of estimates can be. Even in the deliberately narrow scenario studied, the estimated risk ranged over roughly four orders of magnitude, depending on the model used and the assumptions made about the patient population.
What the Failures Looked Like
Some of the patterns are directly relevant to clinical intuition. Across the models evaluated, the statements most often missed were subtle expressions of suicidal ideation (搜索), including ambivalent or conditional language and planning without a clear statement of intent. In other words, the tools were weakest in the same places that demand the most clinical skill.
Practical Implications for Clinicians
Several practical implications follow from this work, even though it remains a research framework rather than a deployable safety rating.
What matters most is which tool a patient uses and in what context, not whether "AI" is safe in the abstract. When a patient mentions AI use, the useful questions are concrete: which tool, how it is used, for what, and in what state of mind. Risk is not a property of just "AI"—it is also a property of a particular tool meeting a particular vulnerability.
Newer or bigger models are usually safer, but that is no guarantee. In general, larger models were better at catching safety concerns, and the smallest, oldest models were clearly the weakest. Some older models could not even respond usefully, and those are best avoided. This raises a concern: a mental health-specific chatbot may be marketed as purpose-built while actually running on an older, weaker model. Yet size is not a 100 percent guarantee—the pattern did not hold uniformly, and a model could be strong at one safety task and weak at another.
Some hidden flaws require more than simple safety patches. Most AI tools run automatic checks meant to catch specific problems, like noticing when a message signals suicidal thoughts. Sometimes a single underlying weakness causes a chatbot to miss several risky situations, leading it to make multiple mistakes. Fixing the filter for one risk helps a little, but the user may encounter harm next if they ask about eating disorders (搜索), PTSD, or medications. The research suggests that what protects the patient is fixing the deeper weakness while also keeping humans in the loop and ready to help.
A Clinical Framework for Patient Conversations
In parallel work, the Division of Digital Psychiatry and the Society of Digital Psychiatry hosted a learning collaborative spanning over a dozen psychiatrists, psychologists, social workers, digital navigators, and therapists across six continents. Their goal: equip participants with tools to have informed conversations with patients already engaging with these technologies.
The collaborative developed a practical assessment framework. When a patient discloses AI chatbot use, clinicians should investigate five domains: dose and pattern (how often, for how long, at what times), content (what topics, whether the patient returns to the same themes), function (what the patient gets from it, how it feels afterward), substitution (whether it is replacing previously helpful coping strategies like group therapy or personal connections), and the chatbot itself (which one, and whether the patient would share a conversation transcript).
The group also discussed how AI might interact differently with specific mental health conditions. For patients with OCD (搜索), seeking reassurance from chatbots can reduce anxiety about obsessions in the moment but often strengthens the reassurance-seeking cycle over time. In eating disorders (搜索), patients might ask chatbots about what they ate or how they look; because LLMs tend to be affirming, they may fail to challenge distorted beliefs, potentially leading to unsafe situations. For depression (搜索) and rumination, chatbots designed to continue conversations may support and reinforce negative thinking patterns. Patients who struggle socially might use chatbots as a replacement for human connection, deepening social withdrawal. In mania, psychosis, or delusional thinking, chatbots can elaborate a line of reasoning without recognizing when it deepens unreal thinking.
Addressing the Training Gap
Recognizing an unmet need for clinician training, the Division of Digital Psychiatry created a free interactive training on approaching patients' AI use. It introduces clinicians to six practical areas: literacy, risks, recognition, documentation, response, and output. Participants learn how LLMs function, why certain structural features create predictable vulnerabilities, how to recognize patterns of use warranting further assessment, how to document AI-related factors when they alter the clinical picture, and how to respond thoughtfully when patients disclose concerning experiences.
One of the greatest risks surrounding AI in mental healthcare may be clinician unpreparedness. Patients deserve clinicians who can approach these conversations with curiosity, humility, and enough knowledge to ask better questions—distinguishing between use that appears helpful, use that deserves monitoring, and use that calls for intervention. AI has become part of many patients' lives, and learning how to talk about it has become just another part of good clinical care.
