The AI Trust Gap in Pharma R&D: Why Scientists Are Hesitating and What Must Change
核心洞察
Pharmaceutical industry AI spending is projected to reach $25 billion by 2030, yet adoption in critical discovery domains like ADME prediction (29%) and generative design (42%) remains low.
Only 22% of researchers describe AI as trustworthy, and nearly half feel undertrained, with scientists citing the need for source citation, peer-reviewed training data, and factual accuracy.
The root cause is fragmented, unvalidated data environments—not the AI models themselves—undermining scientists' ability to defend, reproduce, or trace AI-generated outputs.
The pharmaceutical industry is on track to spend $25 billion on AI by 2030, betting that faster discovery, lower attrition, and sharper decision-making will transform R&D. Yet beneath the surface of this historic investment lies a troubling disconnect: AI adoption drops sharply in the very domains where it matters most. According to Drug Discovery News, adoption rates stand at just 42% for generative design, 40% for biomarker analysis, and a mere 29% for ADME prediction. The bottleneck, industry observers note, is not the models themselves but the fragmented data environment underneath them.
"The gap between what leaders are funding and what scientists are willing to act on is the risk underlying one of the largest technology bets our industry has ever made," writes a contributor from CAS (搜索), a division of the American Chemical Society. "Investment in AI does not guarantee adoption, and if the scientists at the bench don't trust the tools, the promised return on that investment will never materialize."
The trust deficit by the numbers
Elsevier (搜索)'s recently published Researcher of the Future: a Confidence in Research report, which surveyed more than 3,200 academic and corporate researchers globally, quantifies the trust problem. While 58% of researchers now use AI for work—up from 37% in early 2024—only 22% describe AI as trustworthy. Nearly half feel undertrained on how to use AI tools effectively.
The report also reveals what researchers say would increase their confidence: nearly six in ten want AI tools that automatically cite sources, 55% want tools trained on the most up-to-date literature, 55% demand high factual accuracy, and 55% want peer-reviewed content. For corporate researchers, 63% emphasize the importance of keeping input data confidential—a reflection of the proprietary, commercially sensitive nature of drug discovery work.
"That's a demanding but reasonable checklist, and it maps to what good scientific practice looks like," said Mirit Eldor, Managing Director of Life Sciences Solutions at Elsevier (搜索).
Why skepticism is a feature, not a bug
The hesitation among scientists may be misinterpreted by executives as cultural resistance, but the deeper issue is structural. Scientific work carries a built-in accountability standard: every conclusion must be defensible, every result reproducible, and every claim traceable to a source.
"A scientist hesitating to act on an AI-generated answer isn't resisting the technology. They're doing what their training requires: refusing to stake a program decision on something they can't defend, reproduce, or trace," the CAS (搜索) contributor notes. "Skepticism, in that context, isn't the obstacle to AI adoption. It's the standard the technology has to clear."
The trust problem compounds in drug discovery because broad AI models trained on the open web were not designed with the complexity of pharmaceutical data in mind. A single chemical compound may appear under thousands of different names across databases and publications, and AI models trained on that inconsistency inherit the confusion. Outputs that appear authoritative can mask a foundational mismatch between the question being asked and the data the model has to work with.
One widely cited case traces a single AI-introduced error through a chain of downstream impacts, illustrating how quickly one bad output can propagate when the underlying data is not built for the work being done. In drug discovery, where a misidentified target carries serious financial, reputational, and patient consequences, the cost of misplaced trust is exceptionally high.
Where AI is delivering value today
Despite the trust gap, AI is already making meaningful contributions in specific areas. Data synthesis stands out as a leading application, with AI pulling together literature, compound data, reaction dynamics, and clinical results that historically lived in separate repositories and took weeks to consolidate.
One practical example is RARE Hope, which used AI-powered analysis to examine nearly 2,000 existing drugs, employing curated knowledge graphs and biological relationship data to score and prioritize drug candidates based on their interaction with disease-related targets. The approach narrowed a long list of possibilities into candidates for expert review and validation.
Target identification and lead optimization are also areas where AI is accelerating workflows. In target ID, AI helps researchers interrogate disease biology, connect evidence across datasets, and prioritize targets for validation. In hit identification and lead optimization, it can narrow thousands of potential compounds to a workable shortlist, making multiple design-make-test-analyze cycles feasible in ways that were not practical before.
The path forward: curated data and traceability
Addressing the trust gap, experts argue, comes down to three elements: a knowledge base scientists can trust, an AI system that leverages content appropriately, and traceability that lets scientists see exactly where an answer originated before they act on it.
"The integrity of an output cannot exceed the integrity of the data feeding it, and in drug discovery, the difference between curated and uncurated content is the difference between an answer worth acting on and one that has to be re-verified," the CAS (搜索) analysis states.
Eldor emphasizes that making scientific data AI-ready requires treating data preparation as a long-term strategy. "Most scientific data is still built for human consumption. Today, it must also be managed and connected in ways machines can understand," she said. This starts with FAIR principles—making data Findable, Accessible, Interoperable, and Reusable—and requires enrichment via semantic layers that map raw data to standard concepts, terms, and relationships.
Ontologies are essential to this effort, allowing AI to recognize that terms such as "Glucagon-like peptide-1" and "GLP-1" refer to the same entity even when different datasets use different language. Combined with semantic enrichment and knowledge graphs, ontologies create a foundation where data is searchable, linked, and reusable.
The agentic AI horizon and the black box risk
Looking ahead, the emergence of agentic AI—systems that go beyond information retrieval to replicate expert reasoning through complex, multi-step problems—promises to shift AI from an assistive tool to an active research partner. Autonomous agents are also beginning to appear in physical labs through robotics platforms that can run experiments, adjust parameters based on results, and feed findings back into the research pipeline.
The risk, however, is the "black box" problem. "If researchers can't follow how an agent reached a conclusion, then trust and reproducibility suffer," Eldor cautioned. "The answer is guided autonomy: predefined workflows and transparent, documented steps that keep humans in control of the logic while AI does the legwork."
The stakes for the industry
The biggest opportunity AI presents is a fundamental change in the economics of drug discovery. The current model—billions of dollars, over a decade of work, and an 80–90% failure rate—means huge areas of unmet need, particularly rare diseases, struggle to attract investment. AI could alter that equation by identifying non-viable candidates sooner and running more cycles in less time.
The biggest risk, Eldor warns, is overpromising. "Anyone using AI daily knows it makes mistakes—confidently and in ways that are hard to spot. In drug discovery, a misidentified target could cost years."
For biopharma companies seeking value from AI in drug discovery, the message from both the CAS (搜索) analysis and Elsevier (搜索)'s research is consistent: the next phase of investment should not be about bigger models and broader data. It must be about curated data foundations, validated scientific content, and systems that make every output traceable. As the CAS contributor concludes: "When we get those right, AI becomes something scientists can move faster with, not something they have to slow down to verify. That's when we'll start to see the return on our industry's investment."
