Skip to main content
Clinical Trials/NCT07626060
NCT07626060RecruitingNot Applicable

Diagnostic Accuracy of Large Language Models (GPT-4o and Claude) in HEART Score Calculation and 30-Day MACE Prediction in Emergency Department Chest Pain Patients: A Prospective Observational Validation Study Against Three-Expert Consensus

Marmara University Pendik Training and Research Hospital1 site in 1 country690 target enrollmentStarted: June 1, 2026Last updated:
Conditions

Trial Snapshot

Phase
Not Applicable
Status
Recruiting
Sponsor
Enrollment
690
Locations
1
Primary Endpoint
Area Under the ROC Curve (AUC) of GPT-4o and Claude HEART Score for 30-Day MACE Prediction

Study Overview

Brief Summary

This prospective observational diagnostic accuracy study evaluates whether large language models (LLMs) - GPT-4o (OpenAI, gpt-4o-2024-11-20) and Claude (Anthropic, claude-sonnet-4-6) - can accurately calculate HEART scores from unstructured Turkish clinical notes and predict 30-day major adverse cardiac events (MACE) in emergency department patients presenting with non-traumatic chest pain.

The study will enroll 600 consecutive adult patients. For each patient, the same anonymized data (free-text anamnesis, ECG report text, troponin value, and age) will be independently processed by both LLMs via separate API calls with deterministic settings (temperature=0, JSON format). A three-expert consensus HEART score - derived through blinded independent scoring by three emergency medicine physicians with majority-vote adjudication - serves as the reference standard for agreement analysis. Actual 30-day MACE (all-cause death, AMI Type 1/2/4b, unplanned revascularization) determined via national health database and telephone follow-up serves as the outcome for diagnostic accuracy analysis.

A secondary documentation-quality sub-study will quantify how spontaneously Turkish emergency anamnesis notes capture HEART score parameters.

Detailed Description

AI SYSTEM SPECIFICATIONS AND PROMPT PROTOCOL Two distinct large language models (LLMs) will be evaluated as index tests: OpenAI GPT-4o (model string: gpt-4o-2024-11-20) and Anthropic Claude (model string: claude-sonnet-4-6). To ensure reproducibility and eliminate stochastic variation, both models will be accessed via standardized API calls using deterministic parameters (temperature = 0, max_tokens = 500, and strict JSON response format). The exact system prompt layout will be locked prior to initialization, and its integrity will be verified using a SHA-256 cryptographic hash. The models will evaluate each patient record independently in zero-shot isolation, with no cross-contamination or conversational history retention between runs.

REFERENCE STANDARD CONSENSUS PROTOCOL The reference standard consists of a structured consensus HEART score established by three independent emergency medicine physicians (each possessing >=3 years of clinical experience and specific training on HEART score criteria). The physicians will review the anonymized clinical charts while remaining strictly blinded to the LLM outputs and the final 30-day MACE outcomes. For each of the 5 HEART components (scored 0, 1, or 2), a majority vote (2/3 agreement) will determine the final component score. In the event of complete disagreement across all three reviewers on a specific component, a fourth independent adjudicator will resolve the tie.

INDETERMINATE RESULTS MANAGEMENT

In strict compliance with STARD-AI 2025 guidelines, cases with missing or uninterpretable parameters within the free-text clinical notes will be classified into predefined indeterminate tiers:

  1. Complete Cases: 0 indeterminate components (eligible for primary diagnostic accuracy analysis).
  2. Partial Indeterminate: Exactly 1 missing component preventing definitive automatic calculation.
  3. Full Indeterminate: >=2 missing components. The proportion of indeterminate classifications will be quantified for both LLMs and evaluated alongside the routine documentation quality of the charts.

Study Design

Study Type
Observational
Observational Model
Cohort
Time Perspective
Prospective

Eligibility Criteria

Ages
18 Years to — (Adult, Older Adult)
Sex
All
Accepts Healthy Volunteers
No

Inclusion Criteria

  • •Age >=18 years
  • •Chief complaint of non-traumatic chest pain at the emergency department
  • •Written informed consent obtained from the patient or legally authorized representative
  • •Availability for 30-day follow-up (reachable by telephone and/or actively registered in the e-Nabiz national health database)

Exclusion Criteria

  • •Traumatic chest pain etiology
  • •ST-elevation myocardial infarction (STEMI) at presentation requiring immediate reperfusion protocol
  • •Refusal or subsequent withdrawal of informed consent
  • •Inability to complete the mandatory 30-day follow-up period
  • •WITHDRAWAL CRITERIA:
  • •Patient or representative requests data withdrawal after initial consent
  • •Administrative identification of retrospective data entry after enrollment

Outcomes

Primary Outcomes

Area Under the ROC Curve (AUC) of GPT-4o and Claude HEART Score for 30-Day MACE Prediction

Time Frame: 30 days after index emergency department visit

AUC calculated separately for GPT-4o and Claude using the Hanley-McNeil method. MACE is defined as a composite of all-cause death, acute myocardial infarction (Type 1/2/4b), and unplanned revascularization within 30 days. HEART score range is 0-10; a higher score indicates a higher risk of MACE. Analysis will be performed on complete cases only (0 indeterminate components).

Secondary Outcomes

  • Sensitivity and Specificity of GPT-4o and Claude HEART Score at Prespecified Thresholds(30 days after index emergency department visit)
  • Component-Level and Total-Score Agreement (Cohen's Kappa) Between LLMs and Expert Consensus(Baseline (At index emergency department visit))
  • Comparative AUC Difference Between GPT-4o and Claude (DeLong Test)(30 days after index emergency department visit)
  • Proportion of Indeterminate Results for GPT-4o and Claude(Baseline (At index emergency department visit))
  • HEART Parameter Documentation Rate in Routine Turkish Anamnesis Notes(Baseline (At index emergency department visit))
  • Subgroup AUC by Age Group and Sex (Algorithmic Bias Assessment)(30 days after the index emergency department visit)

Investigators

Sponsor
Marmara University Pendik Training and Research Hospital
Sponsor Class
Other
Responsible Party
Principal Investigator
Principal Investigator

Emir Ünal

MD, Assistant Professor

Marmara University Pendik Training and Research Hospital

Study Sites (1)

Loading locations...

Similar Trials