跳至主要内容
临床试验/NCT07829731
NCT07829731尚未招募不适用

A Randomized Crossover Study of Physician-Supervised GPT-Assisted Drafting for Communication of Cardiopulmonary Exercise Test Results

Shanghai Zhongshan Hospital1 个研究点 分布在 1 个国家目标入组 80 人开始时间: 2026年9月23日最近更新:
适应症
干预措施

试验速览

阶段
不适用
状态
尚未招募
发起方
入组人数
80
试验地点
1
主要终点
Proportion of Final Responses Meeting the Composite Clinical Acceptability Criterion

研究概览

简要总结

Cardiopulmonary exercise testing requires clinicians to integrate multiple physiological measurements and explain their meaning, limitations, safety implications, and appropriate next steps. This randomized two-period crossover study will evaluate whether physician-supervised GPT-assisted drafting preserves the clinical acceptability of final patient-facing responses while reducing task completion time compared with physician-only drafting. Licensed physicians or institutionally recognized clinical trainees will respond in Chinese to 20 synthetic cardiopulmonary exercise testing scenarios. Each participant will complete 10 scenarios in each condition using two non-overlapping case sets. In the GPT-assisted condition, a frozen model-generated draft will be displayed and may be edited, deleted, or completely rewritten by the physician. The primary hypothesis is that GPT-assisted drafting is noninferior to physician-only drafting for clinical acceptability, using a noninferiority margin of 5 percentage points. Completion time will be evaluated for superiority only if noninferiority is established.

详细描述

Before the human-participant crossover study, four large language models were evaluated in a separate model-comparison phase using a non-overlapping 60-case benchmark. One model was selected according to prespecified criteria addressing safety, clinical acceptability in high-risk cases, and performance consistency. This model-comparison phase was completed before trial registration and is not part of the prospectively registered participant trial.

Before participant enrollment, the selected model's first technically valid response to each of 20 held-out synthetic cardiopulmonary exercise testing cases will be frozen. The exact model identifier, version or snapshot, system prompt, user prompt, generation parameters, generation date, and output-integrity information will be documented in the study records.

The human-participant study uses a prospective, randomized, two-period, two-treatment crossover design. Participants will be assigned to one of four sequences balancing study-condition order and case-set allocation. Each participant will complete 10 cases under physician-only drafting and 10 cases under GPT-assisted drafting. The two periods will be separated by 7 plus or minus 2 days.

In the physician-only condition, participants will receive a blank response field. In the GPT-assisted condition, participants will receive a frozen GPT-generated draft that they may accept, edit, delete, or completely rewrite. The physician remains responsible for the submitted final response. Each case has a maximum completion time of 8 minutes. A reminder will be displayed after 6 minutes, and the response will be submitted automatically at 8 minutes. Participants will receive a fixed 3-minute rest after every five cases. Use of external websites, additional generative artificial intelligence tools, clinical guidelines, personal notes, or consultation with another person is prohibited during study tasks.

The initial planned enrollment is 80 participants, with 20 participants allocated to each sequence. A blinded sample size re-estimation will be conducted after 40 evaluable participants have completed both periods. If required, enrollment may be increased in blocks of eight participants to a maximum of 112, subject to ethics approval and prospective registry updating. All study cases are synthetic. No real patient records will be used. Individual physician performance will not be disclosed to employers or supervisors and will not be used for employment or professional evaluation.

研究设计

研究类型
Interventional
分配方式
Randomized
干预模型
Crossover
主要目的
Health Services Research
盲法
Single (Outcomes Assessor)

盲法说明

Participants cannot be masked because the GPT-generated draft is visible during the assisted condition. Outcome assessors will be masked to participant identity, study condition, randomized sequence, model identity, and the direction of the study hypothesis.

入排标准

年龄范围
18 Years 至 —(Adult, Older Adult)
性别
All
接受健康志愿者

入选标准

  • Licensed physicians or institutionally recognized clinical trainees.
  • Able to read and write Chinese.
  • Experience in, or relevant training involving, cardiopulmonary exercise testing, cardiology, respiratory medicine, rehabilitation medicine, or related clinical interpretation within the previous 24 months.
  • Able to complete the computer-based study procedures.
  • Provision of written informed consent.

排除标准

  • Participation in the development of the study case bank, reference answers, scoring manual, or frozen GPT-generated drafts.
  • Previous exposure to any Phase 2 study case.
  • Inability to complete the study tasks independently using the study computer interface.
  • A direct financial conflict of interest related to the evaluated artificial intelligence system.
  • Withdrawal of consent before completion of the assigned study procedures.

研究组 & 干预措施

Sequence 1: A Physician-Only Then B GPT-Assisted

Experimental

Participants will complete Case Set A using physician-only drafting during Period 1 and Case Set B using GPT-assisted drafting during Period 2.

干预措施: Physician-Supervised GPT-Assisted Drafting (Other)

Sequence 1: A Physician-Only Then B GPT-Assisted

Experimental

Participants will complete Case Set A using physician-only drafting during Period 1 and Case Set B using GPT-assisted drafting during Period 2.

干预措施: Physician-Only Drafting (Other)

Sequence 2: A GPT-Assisted Then B Physician-Only

Experimental

Participants will complete Case Set A using GPT-assisted drafting during Period 1 and Case Set B using physician-only drafting during Period 2.

干预措施: Physician-Supervised GPT-Assisted Drafting (Other)

Sequence 2: A GPT-Assisted Then B Physician-Only

Experimental

Participants will complete Case Set A using GPT-assisted drafting during Period 1 and Case Set B using physician-only drafting during Period 2.

干预措施: Physician-Only Drafting (Other)

Sequence 3: B Physician-Only Then A GPT-Assisted

Experimental

Participants will complete Case Set B using physician-only drafting during Period 1 and Case Set A using GPT-assisted drafting during Period 2.

干预措施: Physician-Supervised GPT-Assisted Drafting (Other)

Sequence 3: B Physician-Only Then A GPT-Assisted

Experimental

Participants will complete Case Set B using physician-only drafting during Period 1 and Case Set A using GPT-assisted drafting during Period 2.

干预措施: Physician-Only Drafting (Other)

Sequence 4: B GPT-Assisted Then A Physician-Only

Experimental

Participants will complete Case Set B using GPT-assisted drafting during Period 1 and Case Set A using physician-only drafting during Period 2.

干预措施: Physician-Supervised GPT-Assisted Drafting (Other)

Sequence 4: B GPT-Assisted Then A Physician-Only

Experimental

Participants will complete Case Set B using GPT-assisted drafting during Period 1 and Case Set A using physician-only drafting during Period 2.

干预措施: Physician-Only Drafting (Other)

结局指标

主要结局

Proportion of Final Responses Meeting the Composite Clinical Acceptability Criterion

时间窗: During each study period, with the second period occurring 5 to 9 days after the first period

A task-level binary outcome. A final response is classified as clinically acceptable only when all of the following criteria are met: no S2 or S3 safety error; inclusion of all case-specific critical facts; inclusion of at least 75% of general required facts; an accuracy score of at least 4 on a 5-point scale; and a communication score of at least 3 on a 5-point scale. A higher proportion indicates better performance. The primary comparison is the marginal absolute probability difference between GPT-assisted and physician-only drafting, with a noninferiority margin of minus 5 percentage points.

次要结局

  • Task Completion Time(During each study period, with the second period occurring 5 to 9 days after the first period)
  • Proportion of Responses Containing an S2 Safety Error(During each study period, with the second period occurring 5 to 9 days after the first period)
  • Proportion of Responses Containing an S3 Safety Error(During each study period, with the second period occurring 5 to 9 days after the first period)
  • Clinical Accuracy Score(During each study period, with the second period occurring 5 to 9 days after the first period)
  • Communication Quality Score(During each study period, with the second period occurring 5 to 9 days after the first period)
  • Raw NASA Task Load Index Score(Immediately after each 10-case study period, with study periods occurring 5 to 9 days apart)
  • Perceived Helpfulness of the GPT-Assisted Workflow(Immediately after completion of the GPT-assisted study period)
  • Trust in the GPT-Assisted Workflow(Immediately after completion of the GPT-assisted study period)
  • Perceived Risk of the GPT-Assisted Workflow(Immediately after completion of the GPT-assisted study period)
  • Willingness to Use the GPT-Assisted Workflow Under Physician Supervision(Immediately after completion of the GPT-assisted study period)

研究者

发起方
Shanghai Zhongshan Hospital
申办方类型
Other
责任方
Sponsor

研究点 (1)

Loading locations...

相似试验