Generative AI Clinical Decision Support Shows Safety and Documentation Gains but No Significant Reduction in Treatment Failure in Kenyan Primary Care Trial
核心洞察
A cluster-randomized trial across 16 Kenyan primary care clinics found that an LLM-based clinical decision support tool did not significantly reduce 14-day treatment failure (2.2% vs 2.0%; aOR 0.77, 95% CI 0.55–1.08, P=0.13).
The AI intervention significantly improved clinical documentation quality, including appropriate diagnosis (aOR 1.74), comprehensive notes (aOR 1.68), and appropriate treatment plans (aOR 1.71).
No safety concerns were identified, with similar rates of hospitalization and death between groups, and patient satisfaction was identical across arms.
A large-scale pragmatic cluster-randomized trial conducted across 16 primary care facilities in Nairobi and Kiambu counties, Kenya, has found that a generative AI-powered clinical decision support system embedded within an electronic medical record (EMR) did not significantly reduce short-term treatment failure, despite clear improvements in clinical documentation quality and no evidence of safety concerns.
The study, published in Nature Medicine, is one of the first randomized controlled trials worldwide to evaluate whether a large language model (LLM)-based tool can improve patient-level outcomes in real-world clinical practice, rather than merely clinician performance on simulated cases.
Between 22 April and 16 July 2025, 17,626 patients were screened, with 9,702 eligible encounters ultimately analyzed. Clinical officers—mid-level practitioners who deliver much of primary care in Kenya amid a severe physician shortage (approximately 0.3 physicians per 1,000 population)—were randomized at the cluster level to either use the EMR with an integrated AI Consult tool (version 2.0, powered by GPT-4o (搜索)) or to provide standard care with the AI feature disabled.
Primary outcome: no significant difference in treatment failure
The primary endpoint—treatment failure within 14 days, defined as re-presentation with unresolved symptoms, unplanned escalation, or a safety event—occurred in 94 patients (2.0%) in the control arm and 102 patients (2.2%) in the intervention arm. The adjusted odds ratio was 0.77 (95% CI 0.55 to 1.08, P=0.13) in both intention-to-treat and per-protocol analyses. Additional adjustments for encounter-specific variables yielded similar results (aOR 0.72, 95% CI 0.50 to 1.03, P=0.07).
Across the 16 facilities, the pooled odds ratio was 0.76 (95% credible interval 0.50 to 1.12), with between-site heterogeneity described as low (τ=0.22, 95% CrI 0.01 to 0.67). The population-average risk difference corresponded to five fewer treatment failures per 1,000 patients treated (mean risk difference −0.005, 95% CrI −0.013 to 0.001).
Professor Bilal Mateen, senior author and Honorary Professor of Machine Learning for Health at the University of Birmingham and Chief AI Officer at PATH, described the findings as "reassuring but also sobering," noting that "the technology appears safe and clearly improves aspects of clinical decision-making, but translating those gains into measurable patient benefit is much more challenging, particularly in everyday primary care."
Documentation quality significantly improved
Among 2,000 encounters reviewed for documentation quality by an independent expert panel of six Kenyan family physicians blinded to allocation, clinical officers using LLM assistance produced higher-quality notes across all domains. Compared to the control arm, they were more likely to record an appropriate diagnosis (aOR 1.74, 95% CI 1.28 to 2.36, P<0.001), a comprehensive clinical note (aOR 1.68, 95% CI 1.24 to 2.27, P<0.001), and an appropriate treatment plan (aOR 1.71, 95% CI 1.25 to 2.34, P<0.001).
Safety assessment
Of 1,000 LLM outputs associated with red alerts—indicating the system's interpretation of documented information was potentially incorrect or harmful—the expert panel rated 494 (49.4%) as definitely safe and appropriate, and 424 (42.4%) as mostly safe and appropriate. Only 11 outputs (1.1%) were judged unsafe and inappropriate. Clinicians fully adhered to the LLM's advice in 195 encounters (19.5%), partially adhered in 573 (57.3%), and did not adhere in 232 (23.2%).
Across all 16 study sites, 33 serious adverse events (27 hospitalizations and 6 deaths) occurred. Independent review confirmed appropriate management and no causal link to the intervention. Post hoc analysis of the composite outcome of death or hospitalization showed no significant difference (aOR 0.77, 95% CI 0.30 to 1.94, P=0.60).
Cost implications and antibiotic prescribing
The mean per-patient LLM cost in the intervention arm was US$0.04 (95% CI 0.04 to 0.04). While overall antibiotic prescribing rates were similar between arms, antibiotic-related costs were lower in the intervention group (mean difference US$−0.15, 95% CI −0.25 to −0.04). Notably, the direct per-patient savings from reduced antibiotic prescribing exceeded the per-patient cost of running the LLM.
Patient satisfaction and consultation experience
Among 826 patients completing satisfaction surveys, scores were identical at the group level (median 4.0, IQR 4.0–5.0), with no difference in the likelihood of reporting high satisfaction (aOR 1.02, 95% CI 0.70 to 1.49, P>0.9). Median consultation time was 11 minutes in both arms, though a post hoc analysis suggested a slight difference (IQR 7–17 min control, 8–17 min intervention, P=0.031).
Sentinel conditions and diabetes risk identification
No differences were observed in prescribing outcomes for correct antibiotic use or incorrect antimalarial prescribing, nor in the diagnosis or management of hypertension (搜索) or acute malnutrition (搜索) in children. However, individuals were more likely to be identified as being at risk of type 2 diabetes in the control arm than in the intervention arm (aOR 0.88, 95% CI 0.78 to 0.98, P=0.023), suggesting the AI tool may have improved detection of individuals with established diabetes, shifting them from the "at-risk" category into known diagnosis.
Limitations and global relevance
The authors acknowledge several limitations: the trial was conducted within a single private network of urban clinics, potentially limiting generalizability to rural or public-sector settings; the 14-day follow-up may have been too short to capture downstream effects; and the observed event rate was lower than anticipated, limiting statistical power. Post hoc simulations indicated that detecting modest differences in rare clinical events would require sample sizes exceeding 100,000 patients.
Professor Alastair Denniston, co-author and Professor of Regulatory Science and Innovation at the University of Birmingham, emphasized that "AI can be integrated safely into real clinical workflows, without undermining patient trust or clinician autonomy—which is a critical foundation for any future impact."
Professor Richard Riley, Professor of Biostatistics at the University of Birmingham and senior author, added: "Robust trials like this are so important to establish the real impact of using AI in practice. They help set realistic expectations of what AI can actually contribute within existing care pathways, and helps guide where future investment and research effort should be focused."
The study was funded by the Gates Foundation (搜索), sponsored by PATH, and conducted with collaborators from the London School of Hygiene and Tropical Medicine and the KEMRI-Wellcome Trust Research Programme, Kenya.
