跳至主要内容
临床试验/NCT07378358
NCT07378358招募中不适用

Evaluation of AI Large Models for Diagnosis and Treatment in Real-World Cases: Multicenter Retrospective Study

First Affiliated Hospital of Fujian Medical University1 个研究点 分布在 1 个国家目标入组 800 人开始时间: 2026年1月1日最近更新:

试验速览

阶段
不适用
状态
招募中
入组人数
800
试验地点
1
主要终点
Diagnostic Accuracy: Assessed by Top-1 accuracy

研究概览

简要总结

This multicenter retrospective study aims to evaluate the diagnostic and therapeutic performance of three large language models-ChatGPT, Gemini and Deepseek-using 800 archived inpatient medical records from urology departments across four tertiary hospitals. The study will focus on the accuracy and applicability of these models in disease recognition, preliminary diagnosis and treatment recommendation generation, in order to explore their potential value and limitations in supporting clinical decision-making in real-world settings.

研究设计

研究类型
Observational
观察模型
Cohort
时间视角
Retrospective

入排标准

年龄范围
18 Years 至 —(Adult, Older Adult)
性别
All
接受健康志愿者

入选标准

  • The case data is sourced from the four hospitals involved in the study, with complete and authentic diagnosis and treatment records.
  • Patients must be 18 years or older, with no gender restrictions.
  • Complete medical records, including the following core information: patient' s basic information, present illness history, past medical history, physical examination, and auxiliary examinations (including laboratory and imaging tests).
  • A clear discharge diagnosis and treatment plan (including therapeutic measures and follow-up arrangements).
  • Medical records have been archived, with objective and accurate information that has not been altered.
  • The patient or their legal representative has provided informed consent, agreeing to the use of their anonymized medical data for research analysis.

排除标准

  • Medical records with significant missing information, such as key clinical details (present illness history, diagnostic or treatment records, etc.).
  • Cases where the diagnosis or treatment plan is unclear, or where treatment has not been fully completed for an initial diagnosis.
  • Cases where the primary diagnosis is not urological.
  • Cases with major errors or inconsistencies in the records that could affect further assessment.
  • Medical records in special formats or images that are not readable (e.g., handwritten notes, non-standard documentation).
  • Patients who have not signed the informed consent form or who refuse to allow their medical data to be used for research.

结局指标

主要结局

Diagnostic Accuracy: Assessed by Top-1 accuracy

时间窗: Through study completion, an average of 3 months

Top-1: Proportion of cases where the model's first diagnosis matches the true primary diagnosis.

Diagnostic Accuracy: Assessed by Top-3 accuracy

时间窗: Through study completion, an average of 3 months

Top-3: Proportion of cases where the true diagnosis appears in the model's top 3.

Diagnostic Completeness

时间窗: Through study completion, an average of 3 months

Proportion of the model's diagnoses that overlap with all diagnoses (primary and secondary) in the case.

Differential Diagnosis Quality

时间窗: Through study completion, an average of 3 months

Evaluated by experts using a Likert 5-point scale, considering factors like common disease coverage, logical clarity, and specificity

Treatment Plan Quality

时间窗: Through study completion, an average of 3 months

Assesses whether the model's treatment suggestions align with clinical guidelines, scored by experts on completeness, appropriateness, and safety.

Analysis Time

时间窗: Through study completion, an average of 3 months

5.Time taken by the AI model to provide diagnoses and treatment suggestions (in seconds), reflecting real-time capability.

次要结局

未报告次要终点

研究者

申办方类型
Other
责任方
Sponsor

研究点 (1)

Loading locations...

相似试验