跳至主要内容
临床试验/NCT06911645
NCT06911645已完成不适用

Evaluating the Performance of LLMs and Clinicians in Complex Diagnostic Cases: A Randomized Controlled Trial

Stanford University2 个研究点 分布在 1 个国家目标入组 70 人开始时间: 2024年12月16日最近更新:

试验速览

阶段
不适用
状态
已完成
入组人数
70
试验地点
2
主要终点
Diagnostic reasoning

研究概览

简要总结

This study will assess the impact of immediate access to a customized version of GPT-4, a large language model, on performance in case-based diagnostic reasoning tasks. Specifically, it will compare this approach to a two-step process where participants first use traditional diagnostic decision support tools to support their diagnostic reasoning before gaining access to the customized GPT-4 model.

详细描述

Artificial intelligence (AI) technologies, particularly advanced large language models like OpenAI's ChatGPT, have the potential to enhance medical decision-making. While ChatGPT-4 was not specifically designed for medical applications, it has demonstrated promise in various healthcare contexts, including medical note-writing, addressing patient inquiries, and facilitating medical consultations. However, its impact on clinicians' diagnostic reasoning remains largely unknown.

Clinical reasoning is a complex process that involves pattern recognition, knowledge application, and probabilistic reasoning. Integrating AI tools like ChatGPT-4 into physician workflows could help reduce clinician workload and decrease the likelihood of missed diagnoses. However, ChatGPT-4 was neither developed nor validated for diagnostic reasoning, and it may produce misleading information, including plausible but incorrect conclusions that could misguide clinicians. If not used appropriately, it may fail to improve-and could even hinder-clinical decision-making. Therefore, it is essential to study how clinicians use large language models to support clinical reasoning before integrating them into routine patient care.

This study will examine how immediate access to a customized version of ChatGPT-4 impacts performance on case-based diagnostic reasoning tasks, compared to a stepwise approach. In the stepwise approach, participants will first use traditional diagnostic decision support tools to support their case reasoning before interacting with a customized ChatGPT-4 model, at which point they will have the opportunity to revise their initial answers.

Participants will be randomized into different study arms and will respond to diagnostic cases by providing three differential diagnoses, along with supporting and opposing findings for each. They will also identify their top diagnosis and propose next diagnostic steps. Independent reviewers, blinded to treatment assignment, will evaluate their responses.

研究设计

研究类型
Interventional
分配方式
Randomized
干预模型
Parallel
主要目的
Diagnostic
盲法
Single (Outcomes Assessor)

盲法说明

The grading of responses will be performed by assessors blinded to participant identity and treatment assignment.

入排标准

性别
All
接受健康志愿者

入选标准

  • Participants must be licensed physicians and have completed at least post-graduate year 1 (PGY1) of medical training.
  • Training in Internal medicine, family medicine, or emergency medicine.

排除标准

  • Not currently practicing clinically.
  • Participated in one of our previous studies that used the same six diagnostic cases.

结局指标

主要结局

Diagnostic reasoning

时间窗: Through study completion, an average of 6 months

The primary outcome will be the percentage of correct responses per case (range: 0 to 100). For each case, participants will be asked to provide their top three differential diagnoses, along with supporting and opposing findings for each. They will receive 1 point for each plausible diagnosis. Supporting and opposing findings will be graded based on correctness, with 1 point for a partially correct response and 2 points for a completely correct response. Participants will then select their top diagnosis, earning 1 point for a reasonable choice and 2 points for the most accurate diagnosis. Finally, they will list up to three next steps for further patient evaluation, with 1 point awarded for a partially correct response and 2 points for a completely correct response. The primary outcome will be analyzed at the case level, comparing performance between the randomized study groups.

次要结局

  • Prompt frequency(Through study completion, an average of 6 months)
  • Sentiment(Through study completion, an average of 6 months)
  • Time Spent Per Case(Through study completion, an average of 6 months)
  • Participant Perceptions of AI in Clinical Reasoning(Through study completion, an average of 6 months)
  • Customized GPT-4's diagnostic reasoning(Through study completion, an average of 6 months)

研究者

申办方类型
Other
责任方
Principal Investigator
主要研究者

Jonathan Chen

Assistant Professor of Medicine (Biomedical Informatics) and of Biomedical Data Science

Stanford University

研究点 (2)

Loading locations...

相似试验