跳至主要内容
临床试验/NCT07696221
NCT07696221尚未招募不适用

Comparison of Clinical Assessment and Large Language Models in Preoperative Risk Classification: A Retrospective Analysis of ChatGPT, DeepSeek, Gemini, and Claude in ASA Physical Status Classification

Marmara University Pendik Training and Research Hospital0 个研究点目标入组 350 人开始时间: 2026年7月21日最近更新:
适应症

试验速览

阶段
不适用
状态
尚未招募
发起方
入组人数
350
主要终点
Agreement between LLM-assigned and reference-standard ASA-PS class

研究概览

简要总结

The American Society of Anesthesiologists Physical Status (ASA-PS) classification is a cornerstone of preoperative risk assessment, yet interrater variability among clinicians is well documented. Large language models (LLMs) have recently demonstrated expert-level performance in several clinical classification tasks, including ASA-PS assignment.

This retrospective observational study evaluates whether four widely used LLMs - ChatGPT, DeepSeek, Gemini, and Claude - can accurately and consistently assign ASA-PS classes from structured, fully anonymized clinical vignettes derived from real preoperative anesthesia evaluations, using a consensus of senior anesthesiologists as the reference standard.

No patient data will be transmitted to third-party platforms. Clinical information will be converted by the investigators into de-identified structured vignettes containing only age range, sex, body mass index range, presence or absence of systemic diseases, functional capacity, and the major/minor nature of the planned surgery, in full compliance with national data protection legislation (KVKK).

详细描述

Adult patients who underwent preoperative anesthesia evaluation before elective surgery at Marmara University Pendik Training and Research Hospital will be included retrospectively. For each patient, demographic data (age, sex, body mass index), systemic comorbidities (hypertension, diabetes mellitus, coronary artery disease, chronic obstructive pulmonary disease, and others), functional capacity (metabolic equivalents, MET), type of planned surgery (major/minor), and the ASA-PS class assigned by the attending anesthesiologist will be recorded.

Clinical data will be anonymized and converted into structured clinical vignettes by the investigators. Vignettes will contain no identifiers, dates, protocol numbers, or rare diagnostic combinations that could directly or indirectly identify a patient.

Standardization of the LLM assessment process: To ensure independence between assessments, each vignette will be evaluated in a separate, history-free session. A new conversation will be initiated in the relevant model for every patient vignette, thereby eliminating the possibility that the model is influenced by its responses to previous vignettes (context anchoring). The ASA-PS class assigned to one vignette will not be carried over as context into the evaluation of any subsequent vignette. Each vignette will be presented to all four models using an identical, standardized prompt requesting only an ASA-PS class (I-VI) with a brief rationale, in a strictly defined output format. Model outputs will play no role in clinical decision-making. External information retrieval by the models will be disabled, and all queries will be completed within a narrow time window to minimize variability in model versions.

Each vignette will be submitted to each model once (single querying). Consequently, the intra-model test-retest reliability of the LLMs will not be assessed; this is acknowledged as a study limitation, consistent with the probabilistic nature of large language models, which may produce between-session variability in their outputs.

Model versions: The current version of each model available at the time of data collection will be used - ChatGPT (GPT-5.5, OpenAI), Gemini (Gemini 3.5, Google DeepMind), DeepSeek (DeepSeek V4, DeepSeek AI), and Claude (Claude Opus 4.8, Anthropic). These versions reflect the versions current at the time of protocol submission; the most recent stable version of each model accessible during data collection will be used, and the exact version and access date will be recorded. Because publicly available chat interfaces may perform automatic background routing to different model tiers, this is acknowledged as a reproducibility limitation.

研究设计

研究类型
Observational
观察模型
Cohort
时间视角
Retrospective

入排标准

年龄范围
18 Years 至 —(Adult, Older Adult)
性别
All
接受健康志愿者

入选标准

  • Age 18 years or older
  • Planned elective surgery
  • Completed preoperative anesthesia evaluation

排除标准

  • Emergency surgical procedures
  • ASA VI (brain death)
  • Incomplete clinical records

结局指标

主要结局

Agreement between LLM-assigned and reference-standard ASA-PS class

时间窗: Through study completion, an average of 3 months

Quadratic weighted Cohen's kappa between each large language model's ASA-PS assignment (ChatGPT, DeepSeek, Gemini, Claude) and the reference standard defined by consensus of a blinded panel of at least three senior anesthesiologists. Agreement of at least "good" level (weighted kappa ≥ 0.60) is hypothesized.

次要结局

  • Overall classification accuracy of each LLM(Through study completion, an average of 3 months)
  • Subgroup error patterns(Through study completion, an average of 3 months)

研究者

发起方
Marmara University Pendik Training and Research Hospital
申办方类型
Other
责任方
Principal Investigator
主要研究者

dilara gocmen

asistan prof

Marmara University Pendik Training and Research Hospital

相似试验

Large Language Models Versus Anesthesiologists for... | 临床试验