跳至主要内容
临床试验/NCT07328815
NCT07328815已完成不适用

Mitigating Automation Bias in Physician-LLM Diagnostic Reasoning Using Behavioral Nudges

Lahore University of Management Sciences1 个研究点 分布在 1 个国家目标入组 72 人开始时间: 2026年1月17日最近更新:
干预措施

试验速览

阶段
不适用
状态
已完成
入组人数
72
试验地点
1
主要终点
Diagnostic reasoning accuracy score

研究概览

简要总结

The goal of this randomized controlled trial is to evaluate whether behavioral nudges can reduce automation bias, the uncritical acceptance of automated output, in physicians using large language models (LLM) like ChatGPT-5.1 for clinical decision-making.

The main question it aims to answer is: Does a dual-mechanism behavioral nudge intervention (baseline accuracy anchoring plus case-specific color-coded confidence signals) reduce physicians' uncritical acceptance of incorrect LLM recommendations?

Researchers will compare physicians who receive LLM recommendations along with a behavioral nudge to those who receive LLM recommendations without the nudge to assess if the nudge reduces automation bias.

Participants will:

  • Evaluate six clinical vignettes accompanied by LLM-generated recommendations (half containing deliberate, clinically significant errors).
  • Control group: Be able to view LLM recommendations in standard format without the nudge.
  • Treatment group: Be able to view ChatGPT's diagnostic accuracy on standard medical datasets as an initial anchor, then receive color-coded confidence signals alongside each recommendation (e.g., red for low confidence).
  • Have their responses evaluated by blinded reviewers using an expert-developed assessment rubric to detect uncritical acceptance of erroneous information.

详细描述

Automation bias represents a critical challenge in modern clinical practice, particularly as artificial intelligence (AI) tools become increasingly embedded in healthcare workflows. This cognitive phenomenon describes the tendency of clinicians to favor suggestions from automated decision-making systems, even when those suggestions are incorrect. As Large Language Models (LLM) such as ChatGPT-5.1 gain traction in medical settings, their potential to reduce errors and improve efficiency must be weighed against a significant concern: these models lack rigorous medical validation and may amplify existing cognitive biases through incorrect or misleading recommendations.

The emergence of automation bias in medical contexts reflects a complex interplay of environmental and psychological factors. Time constraints in high-volume clinical settings create pressure to accept AI-generated recommendations without adequate scrutiny. Financial incentives that prioritize efficiency over thoroughness may further discourage critical evaluation necessary for sound clinical judgment. Cognitive fatigue during extended shifts diminishes physicians' capacity for sustained analytical thinking. These pressures interact with psychological mechanisms including diffusion of responsibility, overconfidence in technological solutions, and cognitive offloading, collectively creating conditions where uncritical acceptance of AI-generated recommendations becomes more likely.

This randomized controlled trial evaluates the effectiveness of a behavioral nudge intervention designed to mitigate automation bias among medical doctors utilizing LLM-generated diagnostic recommendations. The primary objective is to determine whether this intervention improves diagnostic reasoning performance scores when evaluating clinical vignettes that include deliberately flawed LLM recommendations. Secondary objectives include assessing whether physician experience level, gender, and prior LLM experience moderate the intervention's effectiveness, determining differential effectiveness for vignettes across different confidence signals.

This study employs a single-blind, randomized controlled trial with two parallel arms. Participants will be randomly assigned 1:1 to either the intervention or control arm. To eliminate variability from differences in prompting skills, participants will not interact directly with a live LLM interface. Instead, all participants will use a custom-built web platform displaying clinical vignettes with pre-generated LLM recommendations, ensuring identical LLM-generated content for each vignette.

All participants will evaluate six clinical vignettes during a single, proctored session lasting approximately 75 minutes. Three vignettes will contain deliberately introduced clinical reasoning flaws in the LLM recommendations, while three will contain correct recommendations. Vignettes will be presented in randomized order to prevent pattern detection.

研究设计

研究类型
Interventional
分配方式
Randomized
干预模型
Parallel
主要目的
Diagnostic
盲法
Single (Outcomes Assessor)

盲法说明

Single (Outcomes Assessor)

入排标准

性别
All
接受健康志愿者

入选标准

  • Full or Provisionally Registered Medical Practitioners with the Pakistan Medical and Dental Council (PMDC).
  • Completed Bachelor of Medicine, Bachelor of Surgery (MBBS) Exam. The equivalent degree of MBBS in US and Canada is the Doctor of Medicine (MD).
  • Participants must have completed a structured training program on the use of ChatGPT (or a comparable large language model), totaling at least 10 hours of instruction. The program must include hands-on practice related to LLM's key aspects, specifically prompt engineering and content evaluation.

排除标准

  • Any other Registered Medical Practitioners (Full or Provisional) with PMDC (e.g., professionals with Bachelor of Dental Surgery or BDS).

研究组 & 干预措施

ChatGPT Recommendations without a Behavioral Nudge

No Intervention

Participants will evaluate six clinical vignettes. During the trial, they will have access to clinical recommendations from a specific, commercially available LLM (ChatGPT) in addition to conventional diagnostic resources. LLM recommendations for three vignettes will contain deliberately flawed diagnostic information. The cases will be presented in random order. Participants in this arm will not receive any behavioral nudge.

ChatGPT Recommendations alongside a Behavioral Nudge

Active Comparator

Participants will evaluate six clinical vignettes. During the trial, they will have access to clinical recommendations from a specific, commercially available LLM (ChatGPT) in addition to conventional diagnostic resources. LLM recommendations for three vignettes will contain deliberately flawed diagnostic information and for three vignettes it will contain accurate recommendations). The cases will be presented in random order. Participants in this arm will receive a behavioral nudge embedded in the LLM recommendations interface that presents two synchronized cognitive cues when the LLM panel is expanded: (1) an anchoring cue displaying ChatGPT's baseline diagnostic accuracy on standard medical datasets at the top of the panel to set realistic expectations before cue intervention located immediately below, which shows the LLM recommendations alongside a case-specific color-coded confidence signal.

干预措施: Behavioral Nudge Intervention (Other)

结局指标

主要结局

Diagnostic reasoning accuracy score

时间窗: Assessed at a single time point for each case, during the scheduled diagnostic reasoning evaluation session, which takes place between 0-5 days after participant enrollment.

The primary outcome will be the percent correct for each case, ranging from 0 to 100%, where higher scores indicate better diagnostic performance. For each case, participants will be asked for their three leading diagnoses, findings that support each diagnosis, and findings that oppose each diagnosis. For each plausible diagnosis, participants will receive 1 point. Findings supporting the diagnosis and findings opposing the diagnosis will also be graded based on correctness, with 1 point for each correct response. Participants will then be asked to name their top diagnosis they believe is most likely, earning 9 points for a reasonable response and 18 points for the most accurate response. Finally participants will be asked to name up to 3 next steps to further evaluate the patient with 0.5 point awarded for a partially correct response and 1 point for a completely correct response. The primary outcome will be compared at the case-level between the randomized groups.

次要结局

  • Top choice diagnosis accuracy score(Assessed at a single time point for each case, during the scheduled diagnostic reasoning evaluation session, which takes place between 0-5 days after participant enrollment.)

研究者

申办方类型
Other
责任方
Principal Investigator
主要研究者

Ihsan Ayyub Qazi, PhD

Full Professor, PhD

Lahore University of Management Sciences

研究点 (1)

Loading locations...

相似试验