跳至主要内容
临床试验/NCT07037940
NCT07037940已完成不适用

Comparing OpenEvidence and GPT-4 for Diagnostic Reasoning and Management Decisions: A Randomized Trial

Montefiore Medical Center3 个研究点 分布在 1 个国家目标入组 41 人开始时间: 2025年7月3日最近更新:
干预措施

试验速览

阶段
不适用
状态
已完成
入组人数
41
试验地点
3
主要终点
Clinical Reasoning Performance as determined by Rater Scores

研究概览

简要总结

Clinical decision support tools powered by artificial intelligence (AI) are being rapidly integrated into medical practice. Two leading systems currently available to clinicians are OpenEvidence, which uses retrieval-augmented generation to access medical literature, and GPT-4, a large language model. While both tools show promise, their relative effectiveness in supporting clinical decision-making has not been directly compared. This study aims to evaluate how these tools influence diagnostic reasoning and management decisions among internal medicine physicians.

详细描述

Internal medicine attendings and residents are invited to participate in a study investigating how physicians using a RAG-based LLM (OpenEvidence) perform compared to those using a standard general-purpose LLM (ChatGPT) on both diagnostic reasoning and complex management decisions. As AI tools increasingly enter clinical practice, evidence is needed about which approaches best support physician decision-making. This study will help determine if specialized medical knowledge retrieval systems (OpenEvidence) provide advantages over general AI assistants (ChatGPT) when solving real clinical cases.

Participants will complete one 90-minute Zoom session where clinical cases derived from real, de-identified patient encounters will be solved. Participants will be randomly assigned to use either OpenEvidence or ChatGPT and all responses evaluated by blinded scorers using a validated rubric.

研究设计

研究类型
Interventional
分配方式
Randomized
干预模型
Parallel
主要目的
Other
盲法
Single (Outcomes Assessor)

入排标准

年龄范围
25 Years 至 —(Adult, Older Adult)
性别
All
接受健康志愿者

入选标准

  • Internal medicine residents
  • Internal medicine attending physicians

排除标准

  • Not meeting Inclusion Criteria

研究组 & 干预措施

OpenEvidence

Active Comparator

Participants in this arm will use OpenEvidence as their research tool

干预措施: OpenEvidence (Other)

ChatGPT

Active Comparator

Patients in this arm will use Chat-GPT as their research tool

干预措施: GPT-4 (Other)

结局指标

主要结局

Clinical Reasoning Performance as determined by Rater Scores

时间窗: 15-minutes upon completion of cases, up to approximately 90 minutes total

Clinical reasoning performance will be evaluated based upon the accuracy of the rater scores to responses to the surveys administered. Six blinded, trained independent raters will independently score each participant's response using a validated scoring rubric. Possible response scores can range from 0-100% with higher scores indicating increased clinical reasoning performance. Results for each assessment will be summarized by study arm using basic descriptive statistics and analyzed using mixed-effects models to account for within-subject correlation and between-subject factors.

次要结局

  • Time efficiency(Up to approximately 75 minutes)
  • Decision confidence(15-minutes upon completion of cases, up to approximately 90 minutes total)

研究者

申办方类型
Other
责任方
Sponsor

研究点 (3)

Loading locations...

相似试验