Evaluating ChatGPT-4 as a Decision-Making Support Tool for Surgical Trainees
试验速览
- 阶段
- 不适用
- 状态
- 尚未招募
- 发起方
- 入组人数
- 35
- 试验地点
- 1
- 主要终点
- Proportion of correct responses
研究概览
简要总结
This study aims to assess whether ChatGPT-4 can support surgical trainees in clinical decision-making. By comparing the performance of ChatGPT-4 with junior residents, senior residents, and attending surgeons on standardized clinical scenarios, the study seeks to understand the potential role of large language models in surgical education. The ultimate goal is to evaluate whether ChatGPT-4 can be safely integrated as a supplementary educational tool to aid junior residents in developing critical thinking and surgical judgment.
详细描述
Background:
Artificial Intelligence (AI) is rapidly transforming the medical landscape, offering new possibilities in education, diagnostics, and decision support. In surgery, clinical decision-making is a core competency developed progressively through training. ChatGPT-4, a state-of-the-art large language model developed by OpenAI, has demonstrated competence in handling medical queries and clinical reasoning tasks. However, its performance in complex surgical decision-making compared to human trainees remains largely unexplored.
Objective:
The EDuCATe study aims to evaluate the accuracy and reliability of ChatGPT-4's responses to clinical scenarios involving general surgery cases. Specifically, the study compares the model's performance to that of junior residents, senior residents, and attending surgeons to understand if ChatGPT-4 can serve as a safe and effective educational tool for surgical trainees.
Methods:
研究设计
- 研究类型
- Observational
- 观察模型
- Cohort
- 时间视角
- Prospective
入排标准
- 性别
- All
- 接受健康志愿者
- 否
入选标准
- •Actively enrolled or employed in the general surgery residency or department at the participating institution
- •Willingness to participate and complete all clinical case scenarios
- •Consent to participate in the study
排除标准
- •Incomplete responses
- •Use of external assistance (e.g., internet search, AI tools) when answering scenarios, as self-reported in instructions
结局指标
主要结局
Proportion of correct responses
时间窗: Baseline
Binary outcome (correct vs. incorrect decision)
次要结局
- Comparison of accuracy across experience levels(Baseline)
- Confidence level(Baseline)
- Percentage of use of AI for clinical cases evaluation(Baseline)
研究者
Manuela Mastronardi
Medical Doctor
Ospedali Riuniti Trieste
