跳至主要内容
临床试验/NCT07813754
NCT07813754尚未招募不适用

Teaching Internal Medicine and Family Medicine Residents to Reason With Generative AI: A Multi-Site Randomized Controlled Trial (TEACH-AI)

Beth Israel Deaconess Medical Center4 个研究点 分布在 1 个国家目标入组 200 人开始时间: 2026年9月1日最近更新:
适应症
干预措施

试验速览

阶段
不适用
状态
尚未招募
入组人数
200
试验地点
4
主要终点
Total Score on Expert-Development Rubrics

研究概览

简要总结

The purpose of the TEACH-AI study is to assess whether a brief, structured workshop on artificial intelligence can improve the performance of medicine doctors in training (i.e. residents) in their diagnostic and management reasoning.

In this multi-site randomized controlled trial, internal medicine and family medicine residents are assigned either to receive an in-person workshop on safe, effective LLM use before a standardized AI-assisted assessment, or to complete the same assessment before receiving the workshop. Residents will review clinical cases that are fully synthetic, no protected health information is used, using a password protected LLM interface.

详细描述

Large language models (LLMs) have rapidly entered routine use in medical education and clinical practice. Prior randomized trials have shown that, although LLMs can outperform individual clinicians on some reasoning benchmarks, providing physicians with access to an LLM does not necessarily improve their diagnostic or management performance, and LLMs alone may perform better than physician-plus-LLM teams. These findings suggest that human-AI collaboration is fraught, non-trivial and may require explicit training.

TEACH-AI is a pragmatic, multi-site, two-arm randomized controlled trial embedded within protected residency didactic time. The primary objective is to determine whether a single in-person workshop on the basics of LLMs and best practices of prompting/verification strategies improves residents' performance on an AI-assisted assessment, compared with residents who complete the simulation before receiving the workshop. All participants will ultimately receive the same workshop and the same assessment (either workshop-first vs assessment-first).

The simulation consists of multiple fully synthetic vignettes delivered via a password protected LLM interface. Residents interact freely with the LLM using natural-language prompts and then submit structured final responses regarding aspects such as leading diagnosis, differential, management plan, and/or justification. The platform will record prompts, model outputs, final answers, and timing. Vignette scoring combines correctness of the final diagnosis or management plan. Scoring is conducted by blinded faculty using standardized rubrics and then any discrepancies will be resolved through multiple rounds of discussions.

The trial will enroll up to 200 residents across four ACGME-accredited programs (internal medicine at Beth Israel Deaconess Medical Center, Stanford University, and Cambridge Health Alliance; family medicine at AdventHealth Orlando).

研究设计

研究类型
Interventional
分配方式
Randomized
干预模型
Parallel
主要目的
Diagnostic
盲法
Single (Outcomes Assessor)

盲法说明

The grading of responses will be performed by assessors blinded to participant identity and whether workshop-first or assessment-first assignment.

入排标准

性别
All
接受健康志愿者

入选标准

  • Participants must be licensed physicians and have started at least post-graduate year 1 (PGY1) of medical training.
  • Training in Internal medicine or family medicine or emergency medicine.
  • Able to provide informed consent and complete assessments in English.

排除标准

  • Not currently practicing clinically.
  • Resident directly involved in the design of the study, development of the workshop, or creation/piloting of the assessment vignettes.
  • Resident who declines or withdraws consent.

研究组 & 干预措施

Assessment-First

No Intervention

Participants assigned to this group will be tasked with completing the assessment using a password protected LLM interface first in advance of the workshop. The workshop will be delivered during a scheduled mandatory didactic time slot.

Workshop-first

Experimental

Participants assigned to this group will be tasked with completing the assessment using a password protected LLM interface after they participate in a scheduled mandatory didactic time slot.

干预措施: Workshop-First (Other)

结局指标

主要结局

Total Score on Expert-Development Rubrics

时间窗: Within 24 hours of assessment completion.

The primary outcome will be the number of correct responses for all cases using select questions from expert-developed scoring rubrics. The rubrics were developed using a Delphi consensus process by expert physicians as established by Goh et al (Nature Medicine, 2025; DOI: 10.1038/s41591-024-03456-y). The primary outcome will be analyzed at the case level, comparing performance between the randomized study groups, with a higher number of correct responses indicating a better outcome.

次要结局

  • Time Spent on Management(Within 24 hours of assessment completion.)
  • Prompt Count(Within 24 hours of assessment completion.)
  • Management Reasoning Using Expert-Derived Rubrics(Within 24 hours of assessment completion.)
  • Diagnostic Reasoning(Within 24 hours of assessment completion.)

研究者

申办方类型
Other
责任方
Principal Investigator
主要研究者

Jacob Koshy

Instructor in Medicine

Beth Israel Deaconess Medical Center

研究点 (4)

Loading locations...

相似试验