跳至主要内容
临床试验/NCT07718893
NCT07718893尚未招募不适用

A Randomized Controlled Trial of CITE (Clinical Inference Tethered to Evidence), an Evidence-Grounding Retrieve-and-Verify Layer That Flags Unsupported and Inappropriate Recommendations in AI-Generated Care Plans, Versus AI With Safety Guardrails Alone and Unassisted Care, in Medicaid Primary Care

Waymark0 个研究点目标入组 240 人开始时间: 2026年9月1日最近更新:
适应症

试验速览

阶段
不适用
状态
尚未招募
发起方
入组人数
240
主要终点
Diagnostic accuracy of CITE against clinician adjudication

研究概览

简要总结

This trial evaluates CITE, a retrieve-and-verify layer that audits an AI-generated care plan against a full-text evidence corpus and flags patient-specific codifiable safety hazards to the clinician. The co-primary outcomes are how accurately CITE flags these hazards (sensitivity and specificity versus blinded clinician adjudication) and its clinician alert burden and acceptance, compared with AI care plans using safety guardrails alone and with unassisted clinician care, in Medicaid primary care.

详细描述

Patients are randomized 1:1:1 to (1) unassisted clinician care; (2) AI-generated care plan with safety guardrails; (3) AI-generated care plan with safety guardrails plus CITE. CITE audits the finalized plan against a frozen, versioned evidence corpus and returns physician-facing flags for patient-specific codifiable safety hazards (a recommended drug contraindicated by this patient's diagnosis or laboratory value; a drug-allergy conflict; a dropped high-risk medication; a guideline-indicated therapy omitted for an active diagnosis; a stated quantity refuted by the corpus), each with a verbatim quote and citation; the clinician retains decision authority. Randomization uses a deterministic HMAC permuted-block scheme; outcome assessors are blinded to arm. The co-primary outcomes are (1) the diagnostic accuracy (sensitivity and specificity) of CITE against blinded clinician adjudication, and (2) clinician alert burden (flags per encounter) and acceptance, comparing the CITE arm with the guardrail arm; both are estimable at the enrolled sample size because they do not depend on a rare between-arm event. The unresolved codifiable-hazard rate by arm is reported as a descriptive secondary: codifiable hazards are infrequent, so the trial is not powered for a between-arm efficacy contrast on hazard reduction. A prior trial of a different mechanism (a generic deterministic rule-corpus that surfaced roughly 30 or more flags per encounter and was uninformative) was completed with null results and is registered separately; this trial evaluates a materially different, patient-specific intervention and set of outcomes. Determined exempt by WCG IRB (low risk). Analysis is pre-registered on OSF (https://doi.org/10.17605/OSF.IO/ENXCW).

研究设计

研究类型
Interventional
分配方式
Randomized
干预模型
Parallel
主要目的
Health Services Research
盲法
Single (Outcomes Assessor)

入排标准

年龄范围
18 Years 至 —(Adult, Older Adult)
性别
All
接受健康志愿者

入选标准

  • Age 18 years or older.
  • Medicaid-enrolled and attributed to a participating Waymark primary care site.
  • Primary care encounter that requires clinical reasoning (not administrative-only).
  • English-language clinical documentation.

排除标准

  • Age less than 18 years.
  • Hospice or palliative-care-exclusive care plan.
  • Administrative-only or pharmacy-only encounter that does not surface a clinical decision to the supervising clinician.
  • Encounter where the supervising clinician is the principal investigator.
  • Enrollment in a competing AI-safety study within the prior 90 days.

结局指标

主要结局

Diagnostic accuracy of CITE against clinician adjudication

时间窗: Day 1 (index primary care encounter)

Sensitivity and specificity (with positive and negative predictive values) of the CITE checker for clinically consequential codifiable safety hazards, using blinded clinician adjudication of the plan as the reference standard. Every plan contributes, so the estimate does not depend on a rare between-arm event. Exact-binomial 95% confidence intervals; reported overall and by hazard family.

Clinician alert burden (flags surfaced per encounter)

时间窗: Day 1 (index primary care encounter)

Number of safety flags surfaced to the clinician per encounter in the CITE arm versus the guardrail arm, with clinician acceptance rate. Co-primary usability outcome: a verifier that surfaces an unmanageable number of flags is not deployable regardless of sensitivity (the prior-trial mechanism surfaced a median of about 30 per encounter). Pre-registered acceptability ceiling: median CITE flags per encounter at or below three.

次要结局

  • Clinician action on CITE flags(Day 1 (index primary care encounter))
  • Unresolved codifiable safety-hazard rate by arm (descriptive)(Day 1 (index primary care encounter))
  • Correction of codifiable hazards within 30 days(Up to 30 days after the index encounter)
  • Completed referrals within 30 days(Up to 30 days after the index encounter)
  • Clinical safety composite (exploratory)(Day 1 (index primary care encounter))
  • 30-day acute care utilization (exploratory)(Up to 30 days after the index encounter)

研究者

发起方
Waymark
申办方类型
Industry
责任方
Principal Investigator
主要研究者

Sanjay Basu

Principal Investigator

University of California, San Francisco

相似试验