Improving the Reliability of LLMs as Medical Assistants for the General Public: a Proof of Concept Simulation Trial
试验速览
- 阶段
- 不适用
- 状态
- 已完成
- 发起方
- 入组人数
- 527
- 试验地点
- 2
- 主要终点
- Relevant conditions identification of the 3M-6D education GPT group compared with the GPT group
研究概览
简要总结
This study will evaluate whether three-minute six-dimensions education(3M-6D education) can improve the reliability of large language models as medical assistants for the general public. Participants will be randomly assigned to receive or not receive 3M-6D education and then use ChatGPT, Gemini, or non-AI information resources. The study will assess relevant condition identification, disposition concordance, red-flag identification, and NASA-TLX score.
详细描述
This randomized, controlled, proof-of-concept simulation trial will evaluate whether three-minute six-dimensions education (3M-6D education) can improve the reliability of large language models as medical assistants for the general public.
Eligible participants will be randomly assigned in a 1:1:1:1:1 ratio to one of five study groups: the 3M-6D education GPT group, the GPT group, the 3M-6D education Gemini group, the Gemini group, or the control group. Participants in the 3M-6D education GPT and 3M-6D education Gemini groups will receive approximately three minutes of education before using ChatGPT or Gemini.Each participant will be randomly assigned one of 10 standardized clinical scenarios and complete a simulated counseling task in unrestricted natural language within approximately 10 minutes. The study will assess relevant condition identification, disposition concordance, red-flag identification, and NASA-TLX score.
研究设计
- 研究类型
- Interventional
- 分配方式
- Randomized
- 干预模型
- Parallel
- 主要目的
- Health Services Research
- 盲法
- Single (Outcomes Assessor)
盲法说明
Outcome assessors will be blinded to group assignment when evaluating participants' outcomes. Group information will be removed from the outcomes before assessment.
入排标准
- 年龄范围
- 18 Years 至 —(Adult, Older Adult)
- 性别
- All
- 接受健康志愿者
- 是
入选标准
- •Age 18 years or greater, male or female;
- •Completed primary school or higher education;
- •Able to use a smartphone or computer to complete online interaction;
- •No history of acute ischemic stroke, systemic lupus erythematosus, gastric ulcer, pneumonia, acute cardiac infarction, urinary tract infection, uterine fibroids, diabetes, osteoarthritis, or migraine.
- •Able to understand and comply with study procedures and to provide written informed consent.
排除标准
- •Currently or previously employed as a healthcare worker;
- •Previously received systematic medical training;
- •Currently involved in concurrent research that may interfere with the results of the present trial;
- •The investigator considered that the participant had other conditions that might affect compliance or preclude participation.
研究组 & 干预措施
3M-6D education GPT Group
Participants will first be trained in 3M-6D education, then use ChatGPT to complete a consultation task in unrestricted natural language in approximately 10 minutes.
干预措施: three minutes six dimensions education (Behavioral)
3M-6D education Gemini Group
Participants will first be trained in 3M-6D education, then use Gemini to complete a consultation task in unrestricted natural language in approximately 10 minutes.
干预措施: three minutes six dimensions education (Behavioral)
3M-6D education GPT Group
Participants will first be trained in 3M-6D education, then use ChatGPT to complete a consultation task in unrestricted natural language in approximately 10 minutes.
干预措施: ChatGPT (Other)
3M-6D education Gemini Group
Participants will first be trained in 3M-6D education, then use Gemini to complete a consultation task in unrestricted natural language in approximately 10 minutes.
干预措施: Gemini (Other)
Control group
Participants will use non-AI tools such as internet searches and medical websites to complete a consultation task in unrestricted natural language in approximately 10 minutes.
Gemini Group
Participants will use Gemini to complete a consultation task in unrestricted natural language in approximately 10 minutes.
干预措施: Gemini (Other)
GPT Group
Participants will use ChatGPT to complete a consultation task in unrestricted natural language in approximately 10 minutes.
干预措施: ChatGPT (Other)
结局指标
主要结局
Relevant conditions identification of the 3M-6D education GPT group compared with the GPT group
时间窗: Usually within 1 hour.
Relevant conditions identification is defined as the proportion of participants whose final response includes the expert-defined final diagnosis or a relevant differential diagnosis.
Disposition concordance of the 3M-6D education GPT group compared with the GPT group
时间窗: Usually within 1 hour.
Disposition concordance is defined as the proportion of participants whose final care recommendation matches the expert-defined level. The five levels are self-care, routine outpatient care, urgent outpatient care, emergency department visit, and emergency medical services.
Relevant conditions identification of the 3M-6D education Gemini group compared with the Gemini group
时间窗: Usually within 1 hour.
Disposition concordance of the 3M-6D education Gemini group compared with the Gemini group
时间窗: Usually within 1 hour.
次要结局
- Relevant conditions identification of the 3M-6D education GPT group compared with the control group(Usually within 1 hour.)
- Relevant conditions identification of the 3M-6D education Gemini group compared with the control group(Usually within 1 hour.)
- Disposition concordance of the 3M-6D education GPT group compared with the control group(Usually within 1 hour.)
- Disposition concordance of the 3M-6D education Gemini group compared with the control group(Usually within 1 hour.)
- Red-flag identification in the 3M-6D education GPT group compared with the GPT group(Usually within 1 hour.)
- Red-flag identification in the 3M-6D education GPT group compared with the control group(Usually within 1 hour.)
- Red-flag identification in the 3M-6D education Gemini group compared with the Gemini group(Usually within 1 hour.)
- Red-flag identification in the 3M-6D education Gemini group compared with the control group(Usually within 1 hour.)
- NASA Task Load Index score of the 3M-6D education GPT group compared with the GPT group(Usually within 1 hour.)
- NASA Task Load Index score of the 3M-6D education GPT group compared with the control group(Usually within 1 hour.)
- NASA Task Load Index score of the 3M-6D education Gemini group compared with the Gemini group(Usually within 1 hour.)
- NASA Task Load Index score of the 3M-6D education Gemini group compared with the control group(Usually within 1 hour.)
- Relevant conditions identification of the 3M-6D education GPT group compared with the 3M-6D education Gemini group(Usually within 1 hour.)
- Disposition concordance of the 3M-6D education GPT group compared with the 3M-6D education Gemini group(Usually within 1 hour.)
- Red-flag identification in the 3M-6D education GPT group compared with the 3M-6D education Gemini group(Usually within 1 hour.)
- NASA Task Load Index score of the 3M-6D education GPT group compared with the 3M-6D education Gemini group(Usually within 1 hour.)
研究者
Ji Xunming,MD,PhD
Principal Investigator
Capital Medical University
