跳至主要内容
临床试验/NCT06457269
NCT06457269已完成不适用

Evaluating the Potential of Large Language Models for Respiratory Disease Consultations: A Randomized Crossover Trial

North Sichuan Medical College1 个研究点 分布在 1 个国家目标入组 703 人开始时间: 2023年10月1日最近更新:
适应症

试验速览

阶段
不适用
状态
已完成
发起方
入组人数
703
试验地点
1
主要终点
Expert indicators-Accuracy

研究概览

简要总结

The clinical trial aimes to evaluate multiple large language models in respiratory disease consultations by comparing their performance to that of human doctors across three major medical consultation scenarios.

The main question aims to answer are:

  • How do large language models perform in comparison to human doctors in diagnosing and consulting on respiratory diseases across various clinical scenarios?

In three clinical scenarios including the online query section, the disease diagnosis section and the medical explanation section, research assistants or volunteers will be asked to cross-question all LLMs or real doctors using predefined online questions and their own issues. After each questioning session, a short washout period is implemented to eliminate potential biases.

研究设计

研究类型
Interventional
分配方式
Randomized
干预模型
Crossover
主要目的
Diagnostic
盲法
Quadruple (Participant, Care Provider, Investigator, Outcomes Assessor)

入排标准

性别
All
接受健康志愿者

入选标准

  • Self-reported symptoms of common respiratory diseases, such as cough, chest tightness, fever, and wheezing
  • Ability to engage in LLM dialog operations independently or with minimal peer training
  • A health status deemed suitable for study participation by the pulmonology experts

排除标准

  • 1) Excessively poor health status

结局指标

主要结局

Expert indicators-Accuracy

时间窗: For each participant, starting from the day of random conversation, a maximum participation time of one week will be given. As for subjective expert indicators, the evaluation will be conducted within two months.

Based on the doctors' responses to patients' issues, a 5-point scale will be used for scoring by an expert panel: 5- The responses are completely accurate, addressing all of the patient's questions or diagnosing by identifying the key points of the patient's complaint. 4- The responses are mostly accurate, generally addressing the patient's questions or diagnosing by identifying the key points of the patient's complaint. 3- The responses are moderately accurate, addressing the patient's questions or diagnosing by identifying the key points of the patient's complaint. 2- The responses are rarely accurate, barely addressing the patient's questions or diagnosing by identifying the key points of the patient's complaint. 1- The responses are very inaccurate, not addressing the patient's questions or diagnosing by identifying the key points of the patient's complaint at all.

Expert indicators-Comprehensiveness

时间窗: For each participant, starting from the day of random conversation, a maximum participation time of one week will be given. As for subjective expert indicators, the evaluation will be conducted within two months.

Based on the doctors' responses to patients' issues, a 5-point scale will be used for scoring by an expert panel: 5-The responses are highly comprehensive, addressing various aspects of potential diseases corresponding to the patient's symptoms, providing detailed advice, and offering its own extended interpretations. 4-The responses are mostly comprehensive, covering most aspects of potential common diseases related to the patient's symptoms, and providing fairly detailed advice. 3-The responses are moderately comprehensive, addressing some aspects of potential common diseases related to the patient's symptoms, and offering basic advice. 2-The responses are rarely comprehensive, failing to consider various aspects of potential common diseases related to the patient's symptoms, and providing very limited advice. 1-The responses are not comprehensive at all, overlooking most potential diseases related to the patient's symptoms, and failing to provide any advice.

Empathy indicators

时间窗: For each participant, starting from the day of random conversation, a maximum participation time of one week will be given. As for subjective empathy indicators, the evaluation will be conducted within two months.

Results from CARE scales concerning the doctor-patient relationship, which were completed by patients following each diagnostic session. Specifically, the online query section does not apply the evaluation of CARE scales.

Expert indicators-Ethical compliance

时间窗: For each participant, starting from the day of random conversation, a maximum participation time of one week will be given. As for subjective expert indicators, the evaluation will be conducted within two months.

Based on the doctor's response to the patient's question, an expert panel will review each item in accordance with the Declaration of Helsinki and the International Code of Medical Ethics which aims to determine whether there are any responses or suggestions that could potentially harm the patient or violate ethical guidelines. The findings will be recorded using binary variables: True-The responses are completely ethical. False-When uncertainties exist, the response includes suggestions for the use of controlled medications and some inappropriate or even counterproductive advice.

Expert indicators-Correctness

时间窗: For each participant, starting from the day of random conversation, a maximum participation time of one week will be given. As for subjective expert indicators, the evaluation will be conducted within two months.

Based on the doctors' responses to patients' issues, a 5-point scale will be used for scoring by an expert panel: 5- The responses are completely correct, with no inappropriate or ambiguous statements. 4- The responses are mostly correct, with most statements being appropriate and unambiguous. 3- The responses are generally correct, although there are inappropriate or ambiguous statements, they are acceptable. 2- The responses are partially correct, with few statements being appropriate or unambiguous. 1- The responses are completely incorrect, with nearly all statements being inappropriate and full of ambiguities.

次要结局

  • Regular indicators-Total number of conversations(For each participant, starting from the day of random conversation, a maximum participation time of one week will be given. After the completion of the dialogues, the system will automatically summarize all objective indicators and dialogue information.)
  • Regular indicators-Total conversation cost ($)(For each participant, starting from the day of random conversation, a maximum participation time of one week will be given. After the completion of the dialogues, the system will automatically summarize all objective indicators and dialogue information.)
  • Regular indicators-Total conversation time (min)(For each participant, starting from the day of random conversation, a maximum participation time of one week will be given. After the completion of the dialogues, the system will automatically summarize all objective indicators and dialogue information.)
  • Regular indicators-Follow-up words(For each participant, starting from the day of random conversation, a maximum participation time of one week will be given. After the completion of the dialogues, the system will automatically summarize all objective indicators and dialogue information.)
  • Regular indicators-Number of output statements(For each participant, starting from the day of random conversation, a maximum participation time of one week will be given. After the completion of the dialogues, the system will automatically summarize all objective indicators and dialogue information.)
  • Regular indicators-Total number of questions(For each participant, starting from the day of random conversation, a maximum participation time of one week will be given. After the completion of the dialogues, the system will automatically summarize all objective indicators and dialogue information.)
  • Regular indicators-Number of input statements(For each participant, starting from the day of random conversation, a maximum participation time of one week will be given. After the completion of the dialogues, the system will automatically summarize all objective indicators and dialogue information.)

研究者

发起方
North Sichuan Medical College
申办方类型
Other
责任方
Principal Investigator
主要研究者

Zining Luo

Principal Investigator

North Sichuan Medical College

研究点 (1)

Loading locations...

相似试验

Evaluating the Potential of Large Language Models... | 临床试验