跳至主要内容
临床试验/NCT07725965
NCT07725965进行中(未招募)不适用

Retrospective Validation of Large Language Models (LLM) for the Prognostic Assessment of Clinical Parameters in Emergency Department and Evaluation of the Impact of Automated Anonymisation Methods

University of Cologne1 个研究点 分布在 1 个国家目标入组 100,000 人开始时间: 2026年7月1日最近更新:
适应症

试验速览

阶段
不适用
状态
进行中(未招募)
入组人数
100,000
试验地点
1
主要终点
Diagnostic accuracy of the LLM predictions (AUROC, F1 score) compared with the clinical gold standard

研究概览

简要总结

This retrospective, non-interventional study evaluates the prognostic performance of open-weight Large Language Models (LLMs) in the setting of a German academic emergency department. Using a full census of all consecutive emergency department cases at University Hospital Cologne between 01 January 2023 and 31 December 2025 (approximately 100,000 cases), the study assesses whether LLMs can make reliable prognostic predictions (e.g., hospital admission, imaging, diagnosis, placement) based on the initial history, vital signs, and triage category. In addition, it quantifies how strongly automated anonymization and perturbation procedures affect the models' diagnostic accuracy. This is an Investigator-Initiated Trial (IIT) with no intervention on patients.

详细描述

The study analyzes a retrospective cohort of all emergency department cases at the Central Emergency Department of University Hospital Cologne (01 January 2023 - 31 December 2025). Data originate from the hospital information system (HIS) and are provided in pseudonymized form via the Medical Data Integration Center (MeDIC) of University Hospital Cologne, acting as an independent trusted third party. Extracted data include sociodemographic data (age/year of birth, sex), clinical vital signs (blood pressure, heart rate, respiratory rate, oxygen saturation, temperature, GCS), medical free text (triage records, physician history and admission findings), and process/outcome data serving as the gold standard (ICD-10 diagnoses, imaging performed, admission status, timestamp/length of stay). LLM processing takes place on premise on local compute clusters or in a contractually secured enterprise environment with zero data retention; open-weight models are used. Two arms are compared: Arm A (original data) vs. Arm B (anonymized/perturbed/synthesized data). The primary endpoint is diagnostic accuracy (AUROC, F1 score) against the documented clinical outcome. Working hypotheses: (1) modern LLMs are non-inferior to the human assessment (non-inferiority); (2) modern anonymization procedures reduce model performance by less than 5% (relative performance loss).

研究设计

研究类型
Observational
观察模型
Other
时间视角
Retrospective

入排标准

性别
All
接受健康志愿者

入选标准

  • All consecutive treatment cases at the Central Emergency Department of University Hospital Cologne during the period 01 January 2023 - 31 December 2025 (full census, consecutive inclusion).

排除标准

  • Documented objection to the scientific use of the data pursuant to Art. 21 General data protection Regulation (GDPR).
  • Cases lacking the minimum data required for analysis (triage/history and documented outcome).

研究组 & 干预措施

Analytic Arm A

Original data will be used for the analysis

Analytic Arm B

Anonymized/perturbed data will be used for the analysis

结局指标

主要结局

Diagnostic accuracy of the LLM predictions (AUROC, F1 score) compared with the clinical gold standard

时间窗: From enrollment to the end of retrospective observation period at 1 year

assessment at the level of the individual emergency department encounter).\]

次要结局

  • Relative performance loss of the models between original data (Arm A) and anonymized/perturbed data (Arm B); hypothesis < 5%.(From enrollment to the end of retrospective observation period at 1 year)
  • Sensitivity, specificity, PPV/NPV for binary endpoints and agreement of the triage assessment (Cohen's kappa / Krippendorff's alpha).(From enrollment to the end of retrospective observation period at 1 year)
  • Sensitivity, specificity, positive predictive value(PPV)/negative predictive value (NPV) for binary endpoints and agreement of the triage assessment (Cohen's kappa / Krippendorff's alpha).(From enrollment to the end of retrospective observation period at 1 year)

研究者

申办方类型
Other
责任方
Principal Investigator
主要研究者

Volker Burst

Prof. Dr.

University of Cologne

研究点 (1)

Loading locations...

相似试验