Large Language Models for Chest X-Ray Image Diagnosis and Report Generation: A Stepwise, Multicenter, Early-Stage Clinical Validation Study
试验速览
- 阶段
- 不适用
- 状态
- 已完成
- 发起方
- 入组人数
- 300
- 试验地点
- 3
- 主要终点
- Report Quality Score and agreement score
研究概览
简要总结
The global shortage of radiologists is a pressing issue, particularly in regions with limited medical resources. Against this backdrop, making report generation the core objective of radiology AI systems not only aligns with the practical needs of radiologists but also better serves patient requirements. With the development of multimodal large models, it has become possible to develop automatic report generation systems for medical images. Although ChatGPT 4o has demonstrated certain capabilities in multiple medical sub - fields, as a closed - source system, it has limitations. Its model generation mechanism is opaque, and issues such as hallucination exist. Recently, Deepseek's open - source multimodal large model, Janus - Pro, an "any to any" model, has the advantages of high performance, low cost, and open - source. Nature published three consecutive articles introducing its stunning features. After training and fine - tuning, Janus - Pro shows great potential in medical image diagnosis and report generation. However, currently, the application of Janus - Pro in image diagnosis has not been evaluated. Most existing models are highly versatile but lack optimization for specific domains, and there is a lack of systematic and multi - dimensional evaluation methods to determine the pros and cons of multimodal large models in medical radiology. Based on these current situations, the purpose of our research is to develop and verify the application value of large models dedicated to medical images in image diagnosis and radiology report generation.
研究设计
- 研究类型
- Observational
- 观察模型
- Cohort
- 时间视角
- Prospective
入排标准
- 性别
- All
- 接受健康志愿者
- 是
入选标准
- •Type of examination: standard digital orthopantomogram of the chest during the study period.
- •Completeness of information: Medical records containing basic information, medical history, etc.
- •Informed consent: written consent form signed by the patient or legal representative.
排除标准
- •Anatomical abnormalities: congenital abnormalities of chest development, history of chest surgery, or severe scoliosis that interfere with image interpretation.
- •Poor image quality: chest radiographs with severe artifacts, exposure abnormalities, or incomplete images.
- •Acutely ill: life-threatening and unable to cooperate with the study process.
- •Mental cognitive problems: suffering from severe mental illness or cognitive impairment, unable to understand the study and sign.
- •Pregnant: Fetus is radiosensitive and physiological changes during pregnancy affect the images.
- •Participation in other studies: concurrent participation in other programs that affect the results of this study.
结局指标
主要结局
Report Quality Score and agreement score
时间窗: No more than 1 week
The evaluation assessed the reports generated by the SCP group and the AI-assisted group, using a five-point Likert scale (5 being the best, 1 being the worst) to subjectively evaluate the capability of the large model in generating imaging reports. The RADPEER scoring system was used to assess the agreement between the original reports and the generated reports, including the clinical significance of any discrepancies.
Pairwise preference
时间窗: No more than 1 week
For each evaluator, they must select the better report between the one generated by the large model and the published report. The proportion of cases where three or more evaluators agreed that the AI-assisted group's reports were superior to those of the SCP group was calculated.
Reading Time
时间窗: No more than 1 week
To evaluate the impact of the large model on workflow efficiency, the reading time-defined as the duration from when a radiologist begins reviewing a chest X-ray to the completion of the final radiology report-was measured and automatically recorded by the system.
次要结局
未报告次要终点
