NEJM Perspective Asks Whether AI Can Say "I Don't Know" as LLM Confabulation Rates Reach 99.6% on Fabricated Drug Names
核心洞察
University of Colorado Anschutz (搜索) researchers published an NEJM perspective examining whether large language models (搜索) can express uncertainty, or epistemic humility, in clinical settings.
A companion study testing LLMs against fabricated drug names derived from Pokémon characters found confabulation rates ranging from 2.7% to 99.6% across models.
Researchers argue that safety and performance benchmarks, including the Rx-LLM Benchmark Suite (搜索), are needed before LLMs are integrated into high-stakes pharmacy workflows.
Large language models (搜索) (LLMs) do not appear to express uncertainty by default, and their failure to do so could pose patient safety risks in high-stakes clinical environments such as pharmacy, according to a perspective published in the New England Journal of Medicine (NEJM) by researchers at the University of Colorado Anschutz (搜索) Medical Campus, the Harvard T.H. Chan School of Public Health, and Harvard Medical School.
The perspective, titled "Can AI Say 'I Don't Know'?", examined LLMs' ability to express uncertainty in medicine. The authors argued that more work is needed to help LLMs in clinical workflows achieve what they term epistemic humility, defined in the paper as "a human virtue that involves metacognitive awareness, a moral commitment to truthfulness, and recognition of the limits of one's knowledge." They further emphasized that researchers and clinicians must implement benchmarks to evaluate the efficacy and safety of LLMs in healthcare settings.
Confabulation Rates Up to 99.6% With Fabricated Drug Names
The concern is grounded in empirical findings. In a separate project, the research team evaluated LLM performance in identifying fabricated medications, publishing a paper titled "Drug or Pokémon? Large language model performance in identification of fabricated medications."
The team compiled lists of real medications using both generic and brand names, then added fictitious drugs in the form of Pokémon character names, assigning them a range of plausible doses, routes of administration, and dosing frequencies. Multiple LLMs were evaluated against these datasets, with researchers searching for confabulations—errors in which the AI either missed or overlooked the false drug entirely, or automatically converted the Pokémon character into a different drug.
AI confabulations ranged between 2.7% and 99.6%, with some models performing better than others. Mitigation prompts helped reduce confabulations, but the data suggest that LLMs do not always default to working with drug data in a safe, effective manner.
Why Drug Data Challenges AI Systems
Andrea Sikora, PharmD, MSCR, FCCP, FCCM, an associate professor of biomedical informatics at the University of Colorado Anschutz (搜索) and a clinical pharmacist specializing in critical care, described the structural difficulties drug data present to AI systems.
"Drug data is very high dimensional," Sikora said. "For example, if I told you that I took ibuprofen this morning, there are more questions that would need to be answered. You would need to know the strength, how many tablets I took, whether I took any yesterday, whether I have any cardiovascular issues or renal issues, if I took any other pain medications—all of these factors are relevant to that particular datapoint."
She also pointed to the alphanumeric syntax of drug data—"aspirin 81 milligrams," for instance—as an unusual format for AI systems to parse. Drug naming adds a further complication: most drugs carry multiple names due to pharmaceutical branding, and new drugs enter the market continuously.
Sikora contrasted the pharmacist's trained response to an unfamiliar drug with the behavior of an LLM. "If I saw a drug that I didn't recognize, I wouldn't just assume it was a drug and that the patient was on that drug. I would make it clear that I didn't know what this drug was, and I would ask questions to clarify," she said.
Stakes in Critical Care Pharmacy
Sikora framed the risk in blunt terms. "Drugs can kill you," she said. "Medication errors (搜索) are already a leading cause of death in the United States... [drugs] are a deeply unforgiving thing to make a mistake about. Maybe LLMs are going to make us better, and that certainly would be exciting. I'm excited about this. On the other side, you could say we have now introduced an entirely new way to mess things up. So that's worthy of intrigue or concern."
She called for independent validation of generative AI. "Generative AI needs to be undergoing safety testing from independent validators," Sikora said. "[This is] probably one of the biggest things that I take from this. And that something you are looking for—whether you realize it or not—from your healthcare professional is their humility to realize that you need to be referred to somebody else, or that they don't know. Because, you know, it's your health on the line if they make an assumption."
Benchmarks as a Starting Point
Sikora's research group has developed the Rx-LLM Benchmark Suite (搜索) to test model performance in pharmacy, alongside the Pokémon-based approach. She noted that constructing such tests is not straightforward in areas like pharmacy, where data are high-dimensional and a single clear answer does not always exist. Benchmarking, she argues, is a critical starting point for evaluating the safety and effectiveness of these models before broader clinical integration.
NAM Fellowship to Extend the Work
The research direction has now received national recognition. The National Academy of Medicine (搜索) selected Sikora for its 2026-2028 NAM Fellowship in Pharmacy, part of the NAM Fellowships for Health Sciences Scholars program. She is one of seven fellows in the 2026 class, selected for professional qualifications, reputations, and accomplishments, and will collaborate with researchers, policy experts, and clinicians from across the country during the two-year post.
Sikora joined the Department of Biomedical Informatics faculty in 2024 to strengthen research relevant to clinical pharmacists. Her work focuses on data-driven comprehensive medication management, a framework that uses AI and clinical decision support systems to improve decision-making and clinical care at the patient's bedside. She has received federal grants evaluating AI applications for optimized medication use in the ICU, including a $1.86 million R01 grant from the Agency for Healthcare Research and Quality (搜索) to investigate data-driven optimal pharmacotherapeutic care for critically ill patients and to develop tools that visualize AI-informed predictions for critical care pharmacists tasked with preventing adverse drug events.
The NAM Fellow in Pharmacy also receives a $25,000 flexible research grant. The fellowship is supported through an endowment from the American Association of Colleges of Pharmacy and the American College of Clinical Pharmacy, and allows fellows to participate in the work of an expert study committee or roundtable, engaging with legislators, government officials, and industry leaders.
"So much of my work right now has to do with large language models (搜索) and a big question is how to keep pace and vet this technology to make sure it's safe in clinical settings, so there's a lot of stakeholder involvement to better understand AI's role in clinical pharmacy settings," Sikora said. "Setting safety and performance benchmarks are important in this space, so this fellowship really lends itself to diving deeper and improving technology, processes, and, ultimately, patient care."
