跳至主要内容
临床试验/NCT07298252
NCT07298252已完成不适用

Retrospective Multicenter Multi-Reader, Multi-Case Diagnostic Accuracy Study of Carebot AI MMG Compared With Radiologists on 2D Full-Field Digital Mammography in Breast Cancer Screening

Carebot s.r.o.8 个研究点 分布在 2 个国家目标入组 222 人开始时间: 2025年1月1日最近更新:
干预措施

试验速览

阶段
不适用
状态
已完成
入组人数
222
试验地点
8
主要终点
Balanced accuracy (BA) of Carebot AI MMG for detecting malignant versus non-malignant examinations

研究概览

简要总结

This study evaluates the diagnostic performance of Carebot AI MMG, an artificial intelligence (AI)-enabled medical device for evaluating mammograms. The software analyzes standard full-field digital mammography (FFDM) images and classifies each examination as having no suspicious finding ("Low Risk"), a probably benign mass ("Medium Risk"), or a suspicious malignant mass ("High Risk").

The study is retrospective and observational. It uses anonymized mammography examinations from four screening centers, without any additional imaging or contact with patients. Three experienced breast radiologists independently read the same set of cases, and their assessments are used as the human benchmark. A histopathology-based reference standard, supplemented by radiologist consensus and follow-up information for negative cases, is used to determine whether cancer is present.

The main goal is to compare the AI system with human radiologists in terms of sensitivity and specificity for detecting breast cancer, and to assess whether the AI can achieve non-inferior performance at two predefined operating points: one favoring higher sensitivity and negative predictive value (rule-out) and one favoring higher specificity and positive predictive value (rule-in).

详细描述

Design and setting This is a retrospective, multicenter, multi-reader, multi-case (MRMC) diagnostic accuracy study of Carebot AI MMG, conducted on anonymized 2D full-field digital mammography (FFDM) examinations acquired as part of routine breast cancer screening. Mammograms were collected from four screening centers over a defined time period. No additional imaging was performed for the purpose of this study, and no subjects were contacted.

Data source and population The source dataset consists of 4,729 screening mammography examinations from women aged 32 to 88 years (mean approximately 57 years). Only 2D FFDM studies with a complete set of standard projections (LCC, RCC, LMLO, RMLO) were included. Examinations with incomplete series, unreadable or corrupted DICOM files, or missing/inconsistent key metadata were excluded, as were tomosynthesis (DBT) studies, men, and women under 18 years of age. To ensure sufficient precision of performance estimates, the dataset was enriched with additional biopsy-proven cancers. The final analytical subset comprises 222 examinations, including 48 malignant and 174 non-malignant studies, with representation across three mammography devices (Hologic Selenia Dimensions, Hologic Lorad Selenia, Fujifilm FDR-3000AWS).

Investigational device and comparator The investigational device is Carebot AI MMG (software version 2.9, deep-learning models v2.3), a stand-alone AI system that analyzes 2D FFDM exams and outputs a case-level classification into three categories, together with an internal risk score. Two predefined operating points are evaluated: a high-sensitivity (HSe) threshold, where both benign and malignant masses are treated as "positive" (rule-out setting), and a high-specificity (HSp) threshold, where only malignant masses are counted as "positive" (rule-in setting).

As a human comparator, three experienced radiologists (RAD 1-3) independently read the same anonymized studies using a dedicated DICOM viewer integrated with a labeling application. Radiologists were blinded to AI outputs, clinical information, and outcomes and recorded a case-level classification into the same three categories (Negative/Benign/Malignant). For primary analyses, their binary decisions are derived using the same HSe and HSp rules. In addition, a "random reader" benchmark is constructed for balanced accuracy by repeatedly sampling one radiologist's decision per case in a bootstrap framework (20,000 iterations).

Reference standard The reference standard is established at the study level. A case is labeled "Malignant" if there is histopathological confirmation of breast cancer from biopsy performed in temporal association with the index mammogram. A case is labeled "Non-malignant" if there is consensus between two local radiologists that the finding is negative or stably benign, typically corroborated by at least 2 years of imaging follow-up. Tumor staging (e.g., TNM) is not used in the present analysis. All 48 malignant cases from the participating centers are included in the analytical subset; no cancer-positive examinations were excluded.

研究设计

研究类型
Observational
观察模型
Case Control
时间视角
Retrospective

入排标准

年龄范围
18 Years 至 —(Adult, Older Adult)
性别
Female
接受健康志愿者

入选标准

  • Female sex
  • Age ≥ 18 years at the time of the screening mammogram
  • Screening full-field digital mammography (FFDM) examination with all four standard views (LCC, RCC, LMLO, RMLO) available
  • Sufficient image quality and complete DICOM metadata to allow retrospective analysis

排除标准

  • Age < 18 years
  • Digital breast tomosynthesis (DBT/3D) examinations without a corresponding full 2D FFDM four-view series
  • Incomplete mammography series (missing one or more of LCC, RCC, LMLO, RMLO)
  • Corrupted or unreadable DICOM files
  • Missing or inconsistent key metadata (e.g., laterality, view, acquisition date)

研究组 & 干预措施

Malignant cases

Women with biopsy-proven breast cancer included in the analytical subset (n = 48). Each case corresponds to a screening full-field digital mammography (FFDM) examination with all four standard views (LCC, RCC, LMLO, RMLO), retrospectively identified from participating screening centers.

干预措施: Carebot AI MMG software analysis (Device)

Non-malignant cases

Women without histopathological evidence of breast cancer, classified as negative or stably benign by two independent local radiologists with at least 2 years of imaging follow-up (n = 174). Each case corresponds to a screening FFDM examination with all four standard views (LCC, RCC, LMLO, RMLO), retrospectively selected from the same screening population.

干预措施: Carebot AI MMG software analysis (Device)

结局指标

主要结局

Balanced accuracy (BA) of Carebot AI MMG for detecting malignant versus non-malignant examinations

时间窗: Baseline (index mammography examination; examinations acquired between 01-01-2025 and 14-11-2025; retrospective assessment

Balanced accuracy (BA) is defined as the average of sensitivity and specificity for classifying each mammography examination as malignant or non-malignant. BA will be calculated at two pre-specified operating points of the AI system: a high-sensitivity (HSe) setting and a high-specificity (HSp) setting. Performance will be estimated with 95% confidence intervals and compared to a multi-reader benchmark constructed from three experienced radiologists and a bootstrap-based "random reader" reference.

Sensitivity (Se) of Carebot AI MMG versus histopathology-based reference standard

时间窗: Baseline (index mammography examination; examinations acquired between 01-01-2025 and 14-11-2025; retrospective assessment

Sensitivity is defined as the proportion of malignant examinations correctly classified as positive by the AI system. Sensitivity will be calculated at both the HSe and HSp operating points and compared pairwise with the sensitivity of each of the three radiologists using the same case-level ground truth.

次要结局

  • Specificity (Sp) of Carebot AI MMG versus histopathology-based reference standard(Baseline (index mammography examination; examinations acquired between 01-01-2025 and 14-11-2025; retrospective assessment)
  • Positive predictive value (PPV) of Carebot AI MMG(Baseline (index mammography examination; examinations acquired between 01-01-2025 and 14-11-2025; retrospective assessment)
  • Negative predictive value (NPV) of Carebot AI MMG(Baseline (index mammography examination; examinations acquired between 01-01-2025 and 14-11-2025; retrospective assessment)

研究者

申办方类型
Industry
责任方
Sponsor

研究点 (8)

Loading locations...

相似试验