A Randomized Crossover Study of Physician-Supervised GPT-Assisted Drafting for Communication of Cardiopulmonary Exercise Test Results
试验速览
- 阶段
- 不适用
- 状态
- 尚未招募
- 发起方
- 入组人数
- 80
- 试验地点
- 1
- 主要终点
- Proportion of Final Responses Meeting the Composite Clinical Acceptability Criterion
研究概览
简要总结
Cardiopulmonary exercise testing requires clinicians to integrate multiple physiological measurements and explain their meaning, limitations, safety implications, and appropriate next steps. This randomized two-period crossover study will evaluate whether physician-supervised GPT-assisted drafting preserves the clinical acceptability of final patient-facing responses while reducing task completion time compared with physician-only drafting. Licensed physicians or institutionally recognized clinical trainees will respond in Chinese to 20 synthetic cardiopulmonary exercise testing scenarios. Each participant will complete 10 scenarios in each condition using two non-overlapping case sets. In the GPT-assisted condition, a frozen model-generated draft will be displayed and may be edited, deleted, or completely rewritten by the physician. The primary hypothesis is that GPT-assisted drafting is noninferior to physician-only drafting for clinical acceptability, using a noninferiority margin of 5 percentage points. Completion time will be evaluated for superiority only if noninferiority is established.
详细描述
Before the human-participant crossover study, four large language models were evaluated in a separate model-comparison phase using a non-overlapping 60-case benchmark. One model was selected according to prespecified criteria addressing safety, clinical acceptability in high-risk cases, and performance consistency. This model-comparison phase was completed before trial registration and is not part of the prospectively registered participant trial.
Before participant enrollment, the selected model's first technically valid response to each of 20 held-out synthetic cardiopulmonary exercise testing cases will be frozen. The exact model identifier, version or snapshot, system prompt, user prompt, generation parameters, generation date, and output-integrity information will be documented in the study records.
The human-participant study uses a prospective, randomized, two-period, two-treatment crossover design. Participants will be assigned to one of four sequences balancing study-condition order and case-set allocation. Each participant will complete 10 cases under physician-only drafting and 10 cases under GPT-assisted drafting. The two periods will be separated by 7 plus or minus 2 days.
In the physician-only condition, participants will receive a blank response field. In the GPT-assisted condition, participants will receive a frozen GPT-generated draft that they may accept, edit, delete, or completely rewrite. The physician remains responsible for the submitted final response. Each case has a maximum completion time of 8 minutes. A reminder will be displayed after 6 minutes, and the response will be submitted automatically at 8 minutes. Participants will receive a fixed 3-minute rest after every five cases. Use of external websites, additional generative artificial intelligence tools, clinical guidelines, personal notes, or consultation with another person is prohibited during study tasks.
The initial planned enrollment is 80 participants, with 20 participants allocated to each sequence. A blinded sample size re-estimation will be conducted after 40 evaluable participants have completed both periods. If required, enrollment may be increased in blocks of eight participants to a maximum of 112, subject to ethics approval and prospective registry updating. All study cases are synthetic. No real patient records will be used. Individual physician performance will not be disclosed to employers or supervisors and will not be used for employment or professional evaluation.
研究设计
- 研究类型
- Interventional
- 分配方式
- Randomized
- 干预模型
- Crossover
- 主要目的
- Health Services Research
- 盲法
- Single (Outcomes Assessor)
盲法说明
Participants cannot be masked because the GPT-generated draft is visible during the assisted condition. Outcome assessors will be masked to participant identity, study condition, randomized sequence, model identity, and the direction of the study hypothesis.
入排标准
- 年龄范围
- 18 Years 至 —(Adult, Older Adult)
- 性别
- All
- 接受健康志愿者
- 是
入选标准
- •Licensed physicians or institutionally recognized clinical trainees.
- •Able to read and write Chinese.
- •Experience in, or relevant training involving, cardiopulmonary exercise testing, cardiology, respiratory medicine, rehabilitation medicine, or related clinical interpretation within the previous 24 months.
- •Able to complete the computer-based study procedures.
- •Provision of written informed consent.
排除标准
- •Participation in the development of the study case bank, reference answers, scoring manual, or frozen GPT-generated drafts.
- •Previous exposure to any Phase 2 study case.
- •Inability to complete the study tasks independently using the study computer interface.
- •A direct financial conflict of interest related to the evaluated artificial intelligence system.
- •Withdrawal of consent before completion of the assigned study procedures.
研究组 & 干预措施
Sequence 1: A Physician-Only Then B GPT-Assisted
Participants will complete Case Set A using physician-only drafting during Period 1 and Case Set B using GPT-assisted drafting during Period 2.
干预措施: Physician-Supervised GPT-Assisted Drafting (Other)
Sequence 1: A Physician-Only Then B GPT-Assisted
Participants will complete Case Set A using physician-only drafting during Period 1 and Case Set B using GPT-assisted drafting during Period 2.
干预措施: Physician-Only Drafting (Other)
Sequence 2: A GPT-Assisted Then B Physician-Only
Participants will complete Case Set A using GPT-assisted drafting during Period 1 and Case Set B using physician-only drafting during Period 2.
干预措施: Physician-Supervised GPT-Assisted Drafting (Other)
Sequence 2: A GPT-Assisted Then B Physician-Only
Participants will complete Case Set A using GPT-assisted drafting during Period 1 and Case Set B using physician-only drafting during Period 2.
干预措施: Physician-Only Drafting (Other)
Sequence 3: B Physician-Only Then A GPT-Assisted
Participants will complete Case Set B using physician-only drafting during Period 1 and Case Set A using GPT-assisted drafting during Period 2.
干预措施: Physician-Supervised GPT-Assisted Drafting (Other)
Sequence 3: B Physician-Only Then A GPT-Assisted
Participants will complete Case Set B using physician-only drafting during Period 1 and Case Set A using GPT-assisted drafting during Period 2.
干预措施: Physician-Only Drafting (Other)
Sequence 4: B GPT-Assisted Then A Physician-Only
Participants will complete Case Set B using GPT-assisted drafting during Period 1 and Case Set A using physician-only drafting during Period 2.
干预措施: Physician-Supervised GPT-Assisted Drafting (Other)
Sequence 4: B GPT-Assisted Then A Physician-Only
Participants will complete Case Set B using GPT-assisted drafting during Period 1 and Case Set A using physician-only drafting during Period 2.
干预措施: Physician-Only Drafting (Other)
结局指标
主要结局
Proportion of Final Responses Meeting the Composite Clinical Acceptability Criterion
时间窗: During each study period, with the second period occurring 5 to 9 days after the first period
A task-level binary outcome. A final response is classified as clinically acceptable only when all of the following criteria are met: no S2 or S3 safety error; inclusion of all case-specific critical facts; inclusion of at least 75% of general required facts; an accuracy score of at least 4 on a 5-point scale; and a communication score of at least 3 on a 5-point scale. A higher proportion indicates better performance. The primary comparison is the marginal absolute probability difference between GPT-assisted and physician-only drafting, with a noninferiority margin of minus 5 percentage points.
次要结局
- Task Completion Time(During each study period, with the second period occurring 5 to 9 days after the first period)
- Proportion of Responses Containing an S2 Safety Error(During each study period, with the second period occurring 5 to 9 days after the first period)
- Proportion of Responses Containing an S3 Safety Error(During each study period, with the second period occurring 5 to 9 days after the first period)
- Clinical Accuracy Score(During each study period, with the second period occurring 5 to 9 days after the first period)
- Communication Quality Score(During each study period, with the second period occurring 5 to 9 days after the first period)
- Raw NASA Task Load Index Score(Immediately after each 10-case study period, with study periods occurring 5 to 9 days apart)
- Perceived Helpfulness of the GPT-Assisted Workflow(Immediately after completion of the GPT-assisted study period)
- Trust in the GPT-Assisted Workflow(Immediately after completion of the GPT-assisted study period)
- Perceived Risk of the GPT-Assisted Workflow(Immediately after completion of the GPT-assisted study period)
- Willingness to Use the GPT-Assisted Workflow Under Physician Supervision(Immediately after completion of the GPT-assisted study period)
