跳至主要内容
临床试验/NCT07739121
NCT07739121招募中不适用

Benchmarking AI for Clinical Oncology decisioNmaking (BEACON): A Prospective, Multicentre, Blinded Evaluation of Frontier Large Language Models Against Multidisciplinary Tumour Board Recommendations in Oncology Treatment Planning

Assistance Publique - Hôpitaux de Paris1 个研究点 分布在 1 个国家目标入组 100 人开始时间: 2026年5月1日最近更新:
适应症

试验速览

阶段
不适用
状态
招募中
入组人数
100
试验地点
1
主要终点
Domain-level performance between LLM recommendations and the locked guidelines.

研究概览

简要总结

BEACON (Benchmarking AI for Clinical Oncology decisioNmaking) is a prospective, multicentre, comparative, blinded, non-interventional benchmark evaluating the treatment recommendations of five frontier large language models (LLMs) against the recommendations of multidisciplinary tumour boards (RCP) in oncology treatment planning. One hundred standardised synthetic cases (20 per localisation, across breast, lung, urological, digestive and gynaecological cancers) are submitted as identical structured input to two independent tumour boards per localisation and to five frontier LLMs. Each recommendation - human or model - is decomposed into five predefined decision domains (intent, surgery, radiotherapy, systemic therapy, work-up and biomarkers) and scored 0/1/2 for concordance against a two-tier reference: the consensus of the two tumour boards, complemented by an a priori locked guideline matrix (ESMO, NCCN). The primary endpoint is domain-level concordance between LLM and RCP consensus, expressed as a linearly weighted Cohen's kappa. A co-primary safety endpoint captures the proportion of recommendations carrying serious harm potential, because concordance alone can conceal dangerous errors. Because expert boards may disagree with one another on identical cases, model performance is always interpreted against the human consensus. BEACON is designed as reusable, openly licensed, pre-registered infrastructure: all synthetic cases, evaluation rubrics, the locked guideline matrix, scoring algorithms and verbatim prompts are released for full reproducibility.

详细描述

BEACON is a prospective, multicentre, blinded benchmark using automated, criteria-based scoring. It is built on three design decisions that distinguish it from the existing literature: (i) synthetic, standardised cases remove the record-completeness variability that confounds retrospective comparisons and allow the identical input to be given to every board and every model; (ii) two independent tumour boards per localisation let human-human agreement be measured rather than assumed; and (iii) a guideline matrix, locked a priori, provides an objective anchor applied identically to human and model recommendations.

Reference standard. For each case-domain, a guideline matrix (guideline-recommended / acceptable / unsupported options per case-domain; ESMO, NCCN), locked and time-stamped before data collection, is applied identically to boards and models.

Five decision domains. Every recommendation is decomposed into D1 Intent, D2 Surgery, D3 Radiotherapy, D4 Systemic therapy (class + line), and D5 Work-up & biomarkers before any comparison.

研究设计

研究类型
Observational
观察模型
Cohort
时间视角
Prospective

入排标准

年龄范围
18 Years 至 —(Adult, Older Adult)
性别
All
接受健康志愿者

入选标准

  • Synthetic oncology case within one of the five predefined localisations (breast, lung, urological, digestive, gynaecological).
  • Complete structured schema: UICC 8th-edition stage, biomarkers, ECOG performance status, comorbidities and a standardised clinical question.
  • A clinically answerable treatment-planning question that is mappable to the locked guideline matrix.

排除标准

  • Case outside the five predefined localisations.
  • Incomplete, internally inconsistent or ambiguous schema.
  • Duplicate or near-duplicate of an existing case in the set.
  • Question not resolvable by current guidelines.

结局指标

主要结局

Domain-level performance between LLM recommendations and the locked guidelines.

时间窗: Assessed once at central scoring, after data collection (~October 2026)

For each recommendation domain and each LLM, proportion of LLM recommendation concordant with locked guidelines

次要结局

  • Proportion of recommendations carrying serious harm potential ( LLM and tumour boards)(Up to October 2026)
  • Domain-level recommendation concordance between LLM and tumour-boards(Up to October 2026)
  • Inter-tumour board domain-level recommendation concordance(Up to October 2026)
  • Equipoise rate(Up to October 2026)
  • Completeness(Up to October 2026)
  • Missingness(Up to October 2026)
  • Intensity bias(Up to October 2026)

研究者

申办方类型
Other
责任方
Sponsor

研究点 (1)

Loading locations...

相似试验

Benchmarking Large Language Models Against Tumour... | 临床试验