跳至主要内容
临床试验/NCT07632859
NCT07632859已完成不适用

Diagnostic Accuracy of Two Large Language Models Against a Blinded Specialist Consensus Standard in Turkish Emergency Department Notes: A Retrospective Study of 600 Cases

Marmara University Pendik Training and Research Hospital1 个研究点 分布在 1 个国家实际入组 600 人开始时间: 2026年5月1日最近更新:
适应症

试验速览

阶段
不适用
状态
已完成
发起方
入组人数
600
试验地点
1
主要终点
Diagnostic Accuracy of GPT-4.1 for ICD-10 Chapter-Level Diagnosis

研究概览

简要总结

This retrospective diagnostic accuracy study evaluates two large language models - GPT-4.1 (gpt-4.1-2025-04-14; OpenAI) and Claude Sonnet 4.6 (claude-sonnet-4-6; Anthropic) - as retrospective coding-quality instruments applied to anonymized Turkish-language emergency department anamnesis notes.

The reference standard is the majority consensus of three board-certified emergency medicine specialists who independently coded each note in ICD-10, blinded to one another, to the code entered by the treating physician at case closure, and to the subsequent clinical course. Cases without chapter-level majority agreement are excluded without replacement.

Both models are queried once per note with a single locked prompt at temperature 0 in stateless application programming interface calls, with no retrieval augmentation, no external tools and no extended-reasoning mode. The primary outcome is the proportion of cases in which each model's rank-1 diagnosis matches the reference standard at ICD-10 chapter level, reported with a Wilson 95% confidence interval. Registered secondary outcome measures are chapter-level Cohen's kappa between each model's rank-1 diagnosis and the reference standard; top-3 chapter accuracy for each model; and chapter-level concordance between the closure ICD-10 code and the reference standard. Additional prespecified analyses set out in the statistical analysis plan (paired between-model difference, three-character accuracy, note-length association, confidence calibration and model-to-model agreement) are reported in the primary publication.

The ICD-10 code entered at case closure is characterised against the same reference standard as a description of current documentation practice; it is not a comparator, and no test of superiority or inferiority against model output is performed. The analysis plan was finalised and frozen before any accuracy computation. Reporting follows STARD-AI 2025.

详细描述

STUDY DESIGN: Retrospective diagnostic accuracy study, STARD-AI 2025 reporting, single centre, cohort design.

AI INDEX TESTS: (1) GPT-4.1 (model version gpt-4.1-2025-04-14; OpenAI API). (2) Claude Sonnet 4.6 (model version claude-sonnet-4-6; Anthropic API). Both accessed via the providers' developer application programming interfaces from Python. Temperature = 0. Zero-shot direct prompting with a single locked prompt version; stateless single-turn sessions with no cross-case context, no retrieval augmentation, no external tools and no extended-reasoning mode. No task-specific fine-tuning or additional training was applied; the models were used as released.

MODEL INTERPRETABILITY: Interpretability analyses such as SHAP, Grad-CAM or layer-attribution visualisation are not applicable to this study. Because GPT-4.1 and Claude Sonnet 4.6 are accessed as black-box models through proprietary, closed-source commercial interfaces, internal weights, gradients and attention structures are inaccessible for post-hoc interpretability computation.

REFERENCE STANDARD: Three board-certified emergency medicine specialists independently assess each anonymized note, blinded to one another, to the code entered by the treating physician, and to the subsequent clinical course. The primary diagnosis assigned by at least two of three assessors, reduced to ICD-10 chapter level, constitutes the reference standard. Cases in which all three assessors assign different chapters are excluded without replacement. No joint calibration session was held and no adjudication round was performed; each assessor coded once, according to their own clinical judgement.

DATA PRIVACY: All anamnesis notes are de-identified before processing; direct patient identifiers are removed and no patient name is present in any note at any stage. Each case carries a study-specific sequential number that is not a hospital record number, and no file linking study numbers to patient identities was created or retained. Note text is transmitted to commercial application programming interfaces operated by providers established outside Turkiye; all queries are issued through the providers' developer interfaces in stateless single-turn calls, and no patient identifier is present in any submitted text. De-identified notes are stored in an encrypted, access-restricted database. Conducted in accordance with Turkish Personal Data Protection Law no. 6698.

研究设计

研究类型
观察性
观察模型
队列研究
时间视角
回顾性

入排标准

年龄范围
18 Years 至 —(Adult, Older Adult)
性别
All
接受健康志愿者
否

入选标准

  • Adult patients (aged 18 years and older) presenting to the emergency department, evaluated in the ambulatory (green/yellow triage) area.
  • A free-text electronic anamnesis note entered at presentation in the hospital information system (HBYS). No minimum note length and no "sufficient information for diagnosis" requirement was applied, because such a criterion preferentially retains more readily classifiable cases; note length was treated as a covariate rather than as an eligibility threshold. A note was excluded only if all three of the following were absent: any symptom statement, any duration or onset information, and a non-empty anamnesis field.
  • An ICD-10 code entered by the treating emergency physician at case closure. Cases in which this entry was absent or did not form a valid ICD-10 code were retained in the analysis set and counted in the denominator of the closure-code analyses.

排除标准

  • Notes lacking all three of the following: any symptom statement, any duration or onset information, and a non-empty anamnesis field.
  • Pediatric cases (age under 18 years).
  • Patients critically ill and triaged to high-acuity resuscitation areas (Emergency Severity Index [ESI] level 1).
  • Clinical notes containing residual identifying information that cannot be fully de-identified, preventing compliance with data privacy regulations.
  • Non-independent clinical notes consisting solely of a brief cross-reference to a prior hospital visit without a new history entry.

研究组 & 干预措施

Emergency Department Patient Cohort

Consecutive adult patients (aged 18 years and older) evaluated in the ambulatory (green/yellow triage) area of the emergency department, who had a free-text electronic anamnesis note recorded at presentation and an ICD-10 code entered by the treating physician at case closure. No note-completeness or minimum-length requirement was applied. The closure code is characterised against the reference standard as a description of current documentation practice; it is not a comparator, and no test of superiority or inferiority against model output is performed.

结局指标

主要结局

Diagnostic Accuracy of GPT-4.1 for ICD-10 Chapter-Level Diagnosis

时间窗: At the single index-test run on 3 August 2026

Proportion of cases in which the GPT-4.1 primary (rank 1) diagnosis matches the 3-specialist majority-vote reference standard at the ICD-10 chapter level (22 categories). Range: 0 to 1.00.

Diagnostic Accuracy of Claude Sonnet 4.6 for ICD-10 Chapter-Level Diagnosis

时间窗: At the single index-test run on 3 August 2026

Proportion of cases in which the Claude Sonnet 4.6 primary (rank 1) diagnosis matches the 3-specialist majority-vote reference standard at the ICD-10 chapter level (22 categories). Range: 0 to 1.00.

次要结局

  • Cohen's Kappa Between GPT-4.1 Primary Diagnosis and the Reference Standard(At the single index-test run on 3 August 2026)
  • Cohen's Kappa Between Claude Sonnet 4.6 Primary Diagnosis and the Reference Standard(At the single index-test run on 3 August 2026)
  • Top-3 Diagnostic Accuracy of GPT-4.1(At the single index-test run on 3 August 2026)
  • Top-3 Diagnostic Accuracy of Claude Sonnet 4.6(At the single index-test run on 3 August 2026)
  • Chapter-Level Concordance Between the Closure ICD-10 Code and the Reference Standard(At the original clinical encounter (retrospective data spanning 1 May to 3 August 2026))

研究者

发起方
Marmara University Pendik Training and Research Hospital
申办方类型
其他
责任方
主要研究者
主要研究者

Emir Ünal

MD, Assistant Professor

Marmara University Pendik Training and Research Hospital

研究点 (1)

Loading locations...

标识符

NCT 编号
NCT07632859
其他研究编号
09.2026.26-0514

日期

首次提交
(3个月前)
首次发布
(3个月前)
主要完成日期
(上个月)
研究完成日期
(上个月)
最近核实
(2个月前)
最近更新
(上个月)

监管与共享

FDA 监管药物
否
FDA 监管器械
否
是否有结果
否

相似试验