跳至主要内容
临床试验/NCT07457840
NCT07457840Enrolling By Invitation不适用

Transforming Clinical Decision Support Systems: Using Continuous Bayesian Updates to Integrate AI Predictions With Clinician Expertise

University of California, San Francisco1 个研究点 分布在 1 个国家目标入组 100 人开始时间: 2026年2月15日最近更新:
干预措施

试验速览

阶段
不适用
状态
Enrolling By Invitation
入组人数
100
试验地点
1
主要终点
Clinician Diagnostic Accuracy

研究概览

简要总结

Optimizing the interaction between the human and the machine is a major topic when deploying artificial intelligence (AI) at the bedside. The goal of this randomized clinical vignette study is to learn if presenting AI model outputs via continuous Bayesian updates and/or uncertainty quantification can improve diagnostic accuracy and clinician trust in healthcare professionals (physicians, residents, fellows, physician assistants (PAs), and nurse practitioners (NPs)) from US academic institutions evaluating patients with chest pain or dyspnea.

The main questions it aims to answer are:

  • Does presenting AI predictions as Bayesian-updated post-test probabilities improve diagnostic accuracy compared to standard predicted probabilities?
  • Does the addition of uncertainty quantification (95% confidence intervals) to AI predictions improve diagnostic accuracy?
  • Do these interventions (Bayesian updating and/or uncertainty quantification) help clinicians recover from the negative effects of intentionally misleading AI predictions?

Comparison: Researchers will compare standard AI predicted probabilities (presented without uncertainty) to Bayesian-updated post-test probabilities and/or outputs containing 95% confidence intervals to see if the interventions improve diagnostic accuracy, clinician confidence, and resilience against misleading AI.

Participants will:

  • Review 8 clinical vignettes (simulated patient cases) focusing on chest pain or dyspnea.
  • Provide an initial "pre-test" diagnostic probability for 5 possible diagnoses based on the clinical history alone.
  • View AI model outputs that vary by experimental condition (standard probability vs. Bayesian update, with or without uncertainty intervals, and accurate vs. misleading).
  • Provide an updated "post-test" diagnostic probability for the diagnoses after viewing the AI output.
  • Select and rank diagnostic tests and therapeutic steps for each vignette. Complete a post-survey regarding their trust in the AI, comfort with the data presentation, and demographics.

详细描述

Study Design: This is a 2x2 factorial within-subjects design. The two factors are (1) Bayesian updating via continuous likelihood ratios (CLR) vs. standard predicted probability, and (2) uncertainty quantification (95% confidence intervals) vs. point estimate only. AI prediction accuracy (accurate vs. intentionally misleading) is varied as a within-subjects stratification factor balanced across all 4 conditions, with half of each participant's vignettes receiving accurate predictions and half receiving misleading predictions. AI predictions are simulated (pre-programmed) for experimental control. Vignette order and condition assignment are independently randomized per participant.

Primary Analysis: Diagnostic accuracy is analyzed using a generalized linear mixed model (GLMM) with fixed effects for CLR, Uncertainty, Misleading, and vignette, and a participant random intercept. Pre-specified secondary analyses examine interactions of presentation format with misleading AI.

Sample Size: Simulation-based power analysis (1,000 Monte Carlo iterations per scenario) was conducted using the planned GLMM. Assuming 70% baseline diagnostic accuracy and within-participant ICC of 0.25, the study achieves 85.8% power for the CLR main effect and 85.7% for the Uncertainty main effect with N=100 at alpha=0.05 (two-tailed).

研究设计

研究类型
Interventional
分配方式
Randomized
干预模型
Factorial
主要目的
Other
盲法
Single (Participant)

入排标准

年龄范围
18 Years 至 —(Adult, Older Adult)
性别
All
接受健康志愿者

入选标准

  • Must hold one of the following clinical roles: Nurse Practitioner (NP), Physician Assistant/Physician Associate (PA), Resident Physician, Physician Fellow, or Attending Physician
  • Able to complete the survey in English
  • Access to a computer or tablet (mobile phones are not recommended due to the visual nature of the survey)

排除标准

  • Does not hold an eligible clinical role as defined above
  • Completes fewer than 2 of 8 clinical vignettes (less than 25% of the survey)
  • Has previously participated in this study
  • Unable to complete the survey in English

研究组 & 干预措施

Bayesian Updating (CLR) + Uncertainty (95% CI)

Experimental

AI model prediction is used to perform Bayesian updating of the clinician's pre-test probability into a post-test probability using continuous likelihood ratios (CLR). The post-test probability is presented with a 95% confidence interval. The raw AI predicted probability is not shown to the participant. This arm represents the full intervention combining both candidate approaches. Within this arm, half of vignettes contain accurate AI predictions and half contain intentionally misleading predictions (balanced by design). When AI predictions are misleading, confidence intervals are widened by a factor of 1.5x to simulate greater model uncertainty.

干预措施: Uncertainty Quantification (95% Confidence Interval) (Behavioral)

Standard Probability + No Uncertainty (Control)

Active Comparator

AI model prediction is presented as a standard predicted probability for each possible diagnosis (point estimate only), together with the top 3 clinical features driving the prediction. No confidence interval is shown. This is the control condition representing the most common current approach to presenting AI predictions in clinical settings. Within this arm, half of vignettes contain accurate AI predictions and half contain intentionally misleading predictions (balanced by design).

干预措施: Standard AI Predicted Probability (Behavioral)

Bayesian Updating (CLR) + No Uncertainty

Experimental

AI model prediction is used to perform Bayesian updating of the clinician's pre-test probability into a post-test probability using continuous likelihood ratios (CLR). The post-test probability is presented as a point estimate only, without a confidence interval. The raw AI predicted probability is not shown to the participant. Within this arm, half of vignettes contain accurate AI predictions and half contain intentionally misleading predictions (balanced by design).

干预措施: Bayesian-Updated Post-Test Probability (Behavioral)

Standard Probability + Uncertainty (95% CI)

Experimental

AI model prediction is presented as a standard predicted probability for each possible diagnosis, together with the top 3 clinical features driving the prediction. The predicted probability is accompanied by a 95% confidence interval. Within this arm, half of vignettes contain accurate AI predictions and half contain intentionally misleading predictions (balanced by design). When AI predictions are misleading, confidence intervals are widened by a factor of 1.5x to simulate greater model uncertainty.

干预措施: Standard AI Predicted Probability (Behavioral)

Standard Probability + Uncertainty (95% CI)

Experimental

AI model prediction is presented as a standard predicted probability for each possible diagnosis, together with the top 3 clinical features driving the prediction. The predicted probability is accompanied by a 95% confidence interval. Within this arm, half of vignettes contain accurate AI predictions and half contain intentionally misleading predictions (balanced by design). When AI predictions are misleading, confidence intervals are widened by a factor of 1.5x to simulate greater model uncertainty.

干预措施: Uncertainty Quantification (95% Confidence Interval) (Behavioral)

Bayesian Updating (CLR) + Uncertainty (95% CI)

Experimental

AI model prediction is used to perform Bayesian updating of the clinician's pre-test probability into a post-test probability using continuous likelihood ratios (CLR). The post-test probability is presented with a 95% confidence interval. The raw AI predicted probability is not shown to the participant. This arm represents the full intervention combining both candidate approaches. Within this arm, half of vignettes contain accurate AI predictions and half contain intentionally misleading predictions (balanced by design). When AI predictions are misleading, confidence intervals are widened by a factor of 1.5x to simulate greater model uncertainty.

干预措施: Bayesian-Updated Post-Test Probability (Behavioral)

结局指标

主要结局

Clinician Diagnostic Accuracy

时间窗: Day 1 during survey completion

Proportion of correct diagnostic assessments across all vignettes and experimental conditions. For each vignette, participants rate 5 possible diagnoses on a 0-100% probability scale. The diagnosis assigned the highest probability is considered the participant's final diagnosis. Accuracy is determined by comparing the final diagnosis to the ground truth diagnosis established by expert panel consensus (minimum 4 of 5 board-certified physicians in agreement). Analyzed using a generalized linear mixed model (GLMM) with binary outcome (correct vs. incorrect), fixed effects for CLR, uncertainty quantification, misleading AI, and vignette, and a random intercept for participant.

次要结局

  • Change in Diagnostic Probability Estimates(Day 1 during survey completion)
  • Diagnostic Accuracy Under Misleading AI Predictions(Day 1 during survey completion)
  • Clinician Satisfaction With AI Decision Support (Exploratory)(Day 1 during survey completion)

研究者

申办方类型
Other
责任方
Sponsor

研究点 (1)

Loading locations...

相似试验