Human-AI Teaming Improves Oncology Clinical Trial Prescreening Accuracy Without Slowing Workflow
核心洞察
A randomized trial found that AI-augmented prescreening improved chart-level accuracy to 76.1% versus 71.5% for human-alone review in identifying oncology trial candidates.
The Human+AI approach showed the greatest gains on the hardest criteria: biomarker testing, biomarker interpretation, and tumor staging details.
Efficiency was preserved, with no significant difference in median chart review time between Human+AI (37.4 min) and Human-alone (37.8 min) arms.
Fewer than one in twenty adults with cancer enrolls in a clinical trial, and the labor-intensive prescreening process—where clinical research coordinators (CRCs) manually comb through dense, unstructured health records—is a major barrier. A randomized noninferiority trial published in Nature Communications now provides rigorous evidence that augmenting human reviewers with a neurosymbolic AI language model can improve the accuracy of this prescreening without slowing it down.
The study, led by Ravi B. Parikh and colleagues, compared three approaches using retrospectively collected electronic health records (EHRs) from 355 patients with non-small cell lung cancer (搜索) (NSCLC) or colorectal cancer (搜索) (CrCa) treated at a large community oncology practice: an autonomous AI algorithm, a human CRC working alone, and a CRC utilizing AI augmentation (Human+AI).
AI augmentation delivers measurable accuracy gains
The primary outcome—chart-level accuracy, defined as the proportion of 12 eligibility criteria correctly abstracted compared to a gold-standard set developed by three expert clinician reviewers—was 76.1% in the Human+AI arm versus 71.5% in the Human-alone arm. This difference met the prespecified noninferiority margin and was statistically superior (p = 0.002, Rosenthal's r = 0.165).
The AI-alone arm achieved a mean chart-level accuracy of 59.9%, underscoring that the combination of human expertise with AI assistance outperformed either approach in isolation.
"The Human+AI prescreening approach approximated and improved overall chart abstraction accuracy without compromising efficiency," the authors wrote. "These findings highlight the potential for AI tools to complement human abstraction of eligibility criteria from complex medical documents."
Where AI helped most: biomarkers and staging
Significant criterion-level accuracy improvements favoring Human+AI over Human-alone were observed for seven of the twelve criteria studied. The largest gains emerged in areas where humans typically struggle most—extracting nuanced details from unstructured clinical notes:
- Presence of biomarker testing: 93.2% vs. 84.6% (p < 0.001)
- Interpretation of biomarker result: 91.3% vs. 80.8% (p < 0.001)
- Biomarker results: 79.0% vs. 67.9% (p < 0.001)
- M-staging: 57.0% vs. 43.9% (p < 0.001)
- N-staging: 66.3% vs. 50.5% (p < 0.001)
- T-staging: 71.6% vs. 56.3% (p < 0.001)
- Clinical outcome: 35.9% vs. 23.7% (p = 0.006)
No significant differences were observed for ECOG performance status, medications, cancer type, stage group, or response to therapy.
Efficiency preserved, not yet accelerated
A notable finding was that AI assistance did not reduce chart review time. Median time per chart was 37.4 minutes in the Human+AI arm versus 37.8 minutes in the Human-alone arm—a non-significant difference (p = 0.513). The authors suggest this may reflect a shift in workload: rather than spending time hunting for relevant data points, CRCs devoted their effort to careful review and interpretation of AI-extracted outputs.
"This finding suggests that the workload of clinical research staff may have shifted from the identification of chart elements to the careful review and interpretation of the AI-extracted outputs," the researchers noted, adding that this could help avoid automation bias—overreliance on AI predictions.
Insights into human-AI collaboration dynamics
The study revealed nuanced patterns of automation and confirmation bias. For ECOG performance status, Human+AI performance was lower than Human-alone due to substantial AI inaccuracy, suggesting automation bias. Conversely, for outcome assessment, the AI-alone arm outperformed both collaborative and human-only arms, indicating that confirmation bias may have led reviewers to discount accurate AI suggestions.
Notably, the researchers found no situations where accuracy was lower in the Human+AI arm than in both the AI-alone and Human-alone arms—a finding that contrasts with some radiology studies reporting net negative effects of human-AI collaboration.
Study limitations and future directions
The authors acknowledged four key limitations: the study evaluated a single closed-source AI algorithm, limiting generalizability; subgroup analyses did not examine interactions between document length and complexity; the study was powered for the primary accuracy outcome but not for efficiency or subgroup analyses; and document quality issues such as poorly scanned records may have affected results.
Ongoing work (NCT06561230) is assessing whether AI-enabled prescreening translates to higher and more diverse enrollment in a phase III randomized trial. Future research will also examine how accuracy varies by patient subgroups, including race, to evaluate AI fairness.
"If the bottleneck to trial enrollment is attention rather than eligibility, that is a problem AI is unusually well suited to attack," noted Roupen Odabashian, an oncologist commenting on the study's implications for expanding trial access.
