跳至主要内容
临床试验/NCT07414966
NCT07414966尚未招募不适用

Prospective Evaluation of a Model-Agnostic Meta-Verification Framework (SCOUT) for Scalable Clinical Oversight of Large Language Model Outputs in Coronary Heart Disease Diagnosis: A Multi-Reader, Randomized, Crossover Trial

China National Center for Cardiovascular Diseases0 个研究点目标入组 7 人开始时间: 2026年2月19日最近更新:
干预措施

试验速览

阶段
不适用
状态
尚未招募
发起方
入组人数
7
主要终点
Mean physician review time per case (minutes)

研究概览

简要总结

This prospective, multi-reader, randomized crossover trial evaluates SCOUT (Scalable Clinical Oversight via Uncertainty Triangulation), a model-agnostic meta-verification framework that selectively defers unreliable large language model (LLM) predictions to clinicians by triangulating three orthogonal uncertainty signals: model heterogeneity, stochastic inconsistency, and reasoning critique. The trial assesses whether SCOUT-assisted review can reduce physician review time compared with standard manual review of AI-generated diagnoses while maintaining non-inferior diagnostic accuracy in coronary heart disease (CHD) subtyping.

详细描述

Background: Large language models are increasingly deployed in clinical workflows, yet requiring clinician review of every AI output negates the efficiency gains that motivate their adoption. SCOUT addresses this efficiency-safety paradox through algorithmic meta-verification.

The SCOUT framework triangulates three orthogonal external signals to determine case-level uncertainty: (1) Model Heterogeneity - whether a structurally different auxiliary LLM agrees with the primary model; (2) Stochastic Inconsistency - whether repeated sampling from the same model yields divergent outputs; (3) Reasoning Critique - whether an external checker model identifies logical flaws in the chain-of-thought reasoning.

In this crossover trial, 7 clinicians of varying seniority (2 junior residents, 3 senior residents, 2 attending physicians) each review all 110 cases under both standard manual review and SCOUT-assisted review workflows. The study evaluates workflow efficiency (primary endpoint) and diagnostic accuracy (secondary endpoint).

研究设计

研究类型
Interventional
分配方式
Randomized
干预模型
Crossover
主要目的
Diagnostic
盲法
None

入排标准

年龄范围
18 Years 至 —(Adult, Older Adult)
性别
All
接受健康志愿者

入选标准

  • Board-certified or in-training cardiologists at Fuwai Hospital
  • Spanning three experience strata: junior residents, senior residents, attending physicians

排除标准

  • Clinicians involved in the development or optimization of the SCOUT framework
  • Clinicians involved in the gold-standard adjudication process

研究组 & 干预措施

Control (Standard Manual Review)

Active Comparator

Physicians manually review all cases in the control set (n=54) with access to AI predictions and reasoning. No selective deferral.

干预措施: Standard Manual Review Workflow (Diagnostic Test)

Experimental (SCOUT-Assisted Review)

Experimental

Physicians process the intervention set (n=56) through the SCOUT framework. Low-uncertainty cases are auto-accepted; high-uncertainty cases undergo physician review with full audit trail.

干预措施: SCOUT-Assisted Review Workflow (Diagnostic Test)

结局指标

主要结局

Mean physician review time per case (minutes)

时间窗: Through study completion, an average of 2 hours.

Mean time spent by each clinician reviewing and rendering a diagnostic decision per case under each arm. Measured in minutes.

次要结局

  • Diagnostic accuracy (%)(Through study completion, an average of 2 hours.)
  • Computational Return on Investment (ROI)(Through study completion, an average of 2 hours.)

研究者

发起方
China National Center for Cardiovascular Diseases
申办方类型
Other Gov
责任方
Sponsor

相似试验