Evaluation of the Feasibility and Effectiveness of Generative AI-Assisted Multidisciplinary Decision Support in Medical Intensive Care: A Pilot Randomized Controlled Trial
试验速览
- 阶段
- 不适用
- 状态
- 已完成
- 入组人数
- 15
- 试验地点
- 1
- 主要终点
- Daily Physician Satisfaction Score
研究概览
简要总结
This pilot study evaluated the feasibility and usefulness of generative artificial intelligence (AI) as a clinical decision-support tool for physicians working in a medical intensive care unit. Participating physicians were assigned by work period to either use a generative AI system in addition to usual clinical information resources or to use usual resources without generative AI. The assigned condition was then switched so that participants experienced both approaches. During the AI-assisted periods, physicians used de-identified clinical information and considered the AI-generated responses as reference information. All final clinical decisions remained the responsibility of the treating physicians. The study assessed acceptability, usability, satisfaction, perceived decision support, workload, confidence, and learning experience through repeated questionnaires.
详细描述
This was a single-center, open-label, pilot cluster-randomized crossover study involving physicians working in a medical intensive care unit. Each participating physician was observed during a scheduled one-month rotation in the medical intensive care unit. At the beginning of each monthly rotation, participating physicians were divided into two clusters. The clusters were randomized to begin with either the ChatGPT-assisted condition or the control condition. After approximately two weeks, each cluster crossed over to the alternate condition for the remainder of the one-month rotation. This design allowed participating physicians to experience both study conditions within the same rotation.
During the AI-assisted condition, physicians were encouraged to use ChatGPT (OpenAI) as a reference tool to support clinical information review and decision-making. Only non-identifiable clinical information was permitted to be entered into ChatGPT. Patient names, medical record numbers, contact information, and other information that could directly identify an individual patient were not entered. Physicians summarized clinically relevant information in their own words and considered the responses generated by ChatGPT when planning patient management. The Situation-Background-Assessment-Recommendation framework was recommended as an optional structure for organizing clinical information, but its use was not mandatory. Physicians were otherwise free to formulate their queries and interact with ChatGPT according to their clinical needs. A suggested prompt encouraged ChatGPT to present multiple management options, together with their rationale, potential benefits and risks, relevant supporting evidence, and areas of uncertainty. ChatGPT did not make or implement clinical decisions.
During the control condition, physicians used usual information resources, including discussions with other clinicians, multidisciplinary rounds, consultations, textbooks, clinical practice guidelines, PubMed, and other established clinical reference services, without using ChatGPT or other generative AI tools for study-related clinical decision support.
All diagnostic and treatment decisions were made independently by the treating physicians. Repeated questionnaires assessed satisfaction, decision-making experience, confidence, perceived efficiency, workload, educational value, and other aspects of clinical decision support. At study completion, participants also evaluated usability, satisfaction, perceived learning, reliance on ChatGPT, intention for future use, and the extent to which ChatGPT-generated suggestions were reflected in their clinical plans.
研究设计
- 研究类型
- Interventional
- 分配方式
- Randomized
- 干预模型
- Crossover
- 主要目的
- Health Services Research
- 盲法
- None
入排标准
- 年龄范围
- 19 Years 至 —(Adult, Older Adult)
- 性别
- All
- 接受健康志愿者
- 是
入选标准
- •Age 19 years or older.
- •Physicians, including residents, fellows, and attending physicians, working in the medical intensive care unit at Seoul National University Hospital.
- •Scheduled to work as a primary treating physician for at least 5 days during a planned observation period.
- •Able and willing to provide written informed consent.
排除标准
- •Did not provide written informed consent.
- •Withdrew consent from study participation.
研究组 & 干预措施
ChatGPT-Assisted Condition First, Then Control Condition
Physician clusters used ChatGPT-assisted clinical decision support during the first approximately two weeks of their one-month medical intensive care unit rotation. They then crossed over to the control condition and used usual clinical information resources without generative AI for the remainder of the rotation.
干预措施: Generative AI-Assisted Clinical Decision Support (Other)
ChatGPT-Assisted Condition First, Then Control Condition
Physician clusters used ChatGPT-assisted clinical decision support during the first approximately two weeks of their one-month medical intensive care unit rotation. They then crossed over to the control condition and used usual clinical information resources without generative AI for the remainder of the rotation.
干预措施: Usual Clinical Information Resources (Other)
Control Condition First, Then ChatGPT-Assisted Condition
Physician clusters used usual clinical information resources without generative AI during the first approximately two weeks of their one-month medical intensive care unit rotation. They then crossed over to the ChatGPT-assisted clinical decision-support condition for the remainder of the rotation.
干预措施: Generative AI-Assisted Clinical Decision Support (Other)
Control Condition First, Then ChatGPT-Assisted Condition
Physician clusters used usual clinical information resources without generative AI during the first approximately two weeks of their one-month medical intensive care unit rotation. They then crossed over to the ChatGPT-assisted clinical decision-support condition for the remainder of the rotation.
干预措施: Usual Clinical Information Resources (Other)
结局指标
主要结局
Daily Physician Satisfaction Score
时间窗: At the end of each working day during the one-month medical intensive care unit rotation
The score of 10 questionnaire items assessing physicians' satisfaction with their daily clinical work, modified from Shore and Franks (1986) and Suchman et al. (1993). Each item was rated on a 5-point Likert scale from -2 (strongly disagree) to +2 (strongly agree). Negatively worded items were reverse-scored. The mean score ranges from -2 to +2, with higher scores indicating greater satisfaction.
Daily Clinical Decision-Making Score
时间窗: At the end of each working day during the one-month medical intensive care unit rotation
The score of 6 questionnaire items assessing satisfaction with the clinical decision-making process, perceived decision difficulty, clarity of the preferred treatment, availability of relevant information, and identification of factors affecting the decision. Items were modified from Gedney (1994) and Dolan (1999) and rated on a 5-point Likert scale from -2 (strongly disagree) to +2 (strongly agree). Negatively worded items were reverse-scored. The mean score ranges from -2 to +2, with higher scores indicating a more favorable decision-making experience.
Perceived Quality Score for ChatGPT
时间窗: At the end of the one-month medical intensive care unit rotation, after completion of both crossover periods
The score of 8 questionnaire items assessing the perceived information quality, system quality, and service quality of ChatGPT, modified from Pillong et al. (2025). Each item was rated on a 5-point Likert scale from -2 (strongly disagree) to +2 (strongly agree). The negatively worded response-time item was reverse-scored. The mean score ranges from -2 to +2, with higher scores indicating better perceived quality.
Generative AI Usability Score
时间窗: At the end of the one-month medical intensive care unit rotation, after completion of both crossover periods
The score of 3 questionnaire items assessing ease of use, ease of learning, and clarity of interaction with Generative AI (ChatGPT), modified from Pillong et al. (2025). Each item was rated on a 5-point Likert scale from -2 (strongly disagree) to +2 (strongly agree). The mean score ranges from -2 to +2, with higher scores indicating greater usability.
Satisfaction Score for Generative AI Use
时间窗: At the end of the one-month medical intensive care unit rotation, after completion of both crossover periods
The score of 6 questionnaire items assessing the perceived usefulness, productivity, effectiveness, overall satisfaction, appropriateness, and intention to reuse Generative AI (ChatGPT), modified from Pillong et al. (2025). Each item was rated on a 5-point Likert scale from -2 (strongly disagree) to +2 (strongly agree). The mean score ranges from -2 to +2, with higher scores indicating greater satisfaction.
次要结局
未报告次要终点
