跳至主要内容
临床试验/NCT07470463
NCT07470463招募中不适用

AI Medical Interviewing and Diagnostic System Performance Evaluation: One-Shot Vision Differential Diagnosis (OSVDE) and Multi-Step Conversational Non-Inferiority (MSCNE) Evaluation.

Magic Health Inc. (d.b.a. Nolla Health)1 个研究点 分布在 1 个国家目标入组 30 人开始时间: 2026年3月11日最近更新:
干预措施

试验速览

阶段
不适用
状态
招募中
发起方
入组人数
30
试验地点
1
主要终点
Top-1 Diagnostic Accuracy

研究概览

简要总结

This study evaluates the diagnostic performance of a multimodal artificial intelligence (AI) system (AIMD.1) using de-identified medical images and semi-synthetic patient simulations. The study combines retrospective analysis of existing publicly available image datasets with prospective data collection from board-certified clinicians who complete diagnostic evaluation tasks.

In the One-Shot Vision Differential Evaluation (OSVDE) stage, clinicians review individual de-identified medical images and generate a ranked list of potential diagnoses based solely on visual features. In the Multi-Step Conversational Non-Inferiority Evaluation (MSCNE) stage, clinicians complete diagnostic assessments using semi-synthetic patient simulations derived from de-identified medical images. Clinician performance will be compared with the AI system on the same diagnostic tasks.

Human participants consist solely of board-certified clinicians who provide diagnostic responses. Medical images and simulated cases are study materials and are not considered study participants. No identifiable patient data are used, and the AI system is evaluated in an offline research environment and is not used for clinical decision-making or patient care.

详细描述

Artificial intelligence (AI) systems have demonstrated promising capabilities in medical diagnosis; however, rigorous benchmark evaluation is necessary prior to clinical deployment. AIMD.1 is a multimodal AI diagnostic system designed to assist with clinical reasoning through analysis of medical images and conversational diagnostic interactions.

This study is a benchmark performance evaluation of AIMD.1, a multimodal AI medical diagnostic system developed by Nolla Health and designed to assist clinical reasoning through analysis of medical images and structured conversational diagnostic interactions. The evaluation is conducted entirely in an offline research environment; the AI system is not used to guide real-world clinical care or patient management during this study. The study is designed as a benchmark performance evaluation prior to any prospective validation involving real patients.

The global healthcare system faces a significant workforce shortage, with projections suggesting a deficit of up to 11 million practitioners by 2030. AI systems for medical diagnosis have shown promise in addressing this gap in controlled research settings, but rigorous benchmark validation against clinician-level performance is needed before clinical deployment. This study addresses that need by evaluating AIMD.1 against both established AI benchmark systems and directly against board-certified clinicians completing the same diagnostic tasks.

The study employs two complementary evaluation stages designed to assess distinct aspects of diagnostic capability:

Stage 1 - One-Shot Vision Differential Evaluation (OSVDE): The AI system and clinician participants independently review individual de-identified medical images and generate ranked top-5 differential diagnoses based solely on visual features. The image corpus comprises approximately 11,500-15,000 de-identified images spanning at least 12 medical specialties (Dermatology, Internal Medicine, Otolaryngology, Gynecology, Orthopedics, Pediatrics, Geriatrics, Emergency Medicine, Ophthalmology, Endocrinology, Family Medicine, and others) and 48 disease clusters. Image sources include approximately 14,000 images retrieved from standard search engines (Google and Bing) and open access repositories such as the PMC Open Access Dataset, filtered for Creative Commons and public domain licensing, as well as approximately 1,000 de-identified clinical images provided by Nolla Health under terms of service permitting de-identified use for research purposes, with all images de-identified per HIPAA Safe Harbor standards. All source images undergo random combinations of affine and non-affine transformations (blurring, sharpening, contrast adjustment, color adjustment, pixel shifting, rotation, stretching, Gaussian noise, among others) to produce fundamentally distinct images from the originals while preserving clinically relevant visual features on average. Each transformed image is verified by at least one board-certified dermatologist or primary care clinician and relabeled as necessary, with ambiguous or non-clinically relevant images removed from the corpus. This preprocessing pipeline also provides additional de-identification through cropping or masking of potentially identifiable regions, including facial features. Images are localized using disease-name keywords defined in the study's disease ontology and downloaded in standard formats (JPEG, PNG) with associated metadata including ground-truth diagnostic labels and disease categories. Where available, additional metadata such as Fitzpatrick skin type (I-VI) and patient age range (pediatric, adult, geriatric) are recorded to enable subgroup analyses of diagnostic performance across demographic categories.

研究设计

研究类型
Observational
观察模型
Case Only
时间视角
Prospective

入排标准

年龄范围
18 Years 至 —(Adult, Older Adult)
性别
All
接受健康志愿者

入选标准

  • Active board certification in Dermatology, Internal Medicine, Otolaryngology, Gynecology, Orthopedics, Pediatrics, Geriatrics, Emergency Medicine, Ophthalmology, Psychiatry, Endocrinology, Family Medicine, or a closely related specialty
  • Age 18 years or older
  • Ability to complete diagnostic evaluation sessions remotely using a computer or tablet with reliable internet access

排除标准

  • Loss of active board certification in an eligible specialty
  • Inability to complete the evaluation session remotely

研究组 & 干预措施

Clinician Participants

Board-certified clinicians who participate in diagnostic evaluation tasks using de-identified medical images and semi-synthetic patient simulations to assess diagnostic accuracy. Clinicians provide differential diagnoses for benchmark comparison with an AI diagnostic system.

干预措施: AI Diagnostic System (AIMD.1) (Diagnostic Test)

结局指标

主要结局

Top-1 Diagnostic Accuracy

时间窗: At completion of diagnostic evaluations (up to 6 months)

Proportion of evaluated cases in which the primary diagnosis generated by the AI Diagnostic System matches the reference (ground truth) diagnosis. Accuracy will be calculated across de-identified medical image cases and semi-synthetic patient simulation cases and compared with clinician performance.

次要结局

  • Top-5 Diagnostic Accuracy(At completion of diagnostic evaluations (up to 6 months))
  • Diagnostic Accuracy of Clinician Participants(At completion of diagnostic evaluations (up to 6 months))
  • Non-Inferiority of AI Diagnostic Accuracy Compared to Clinicians(At completion of diagnostic evaluations (up to 6 months))
  • Calibration of AI Diagnostic Confidence(At completion of diagnostic evaluations (up to 6 months))
  • Area Under the Receiver Operating Characteristic Curve (AUC) and Precision Recall Curve (PRC)(At completion of diagnostic evaluations (up to 6 months))
  • Time-to-Diagnosis in Conversational Simulations(At completion of simulation evaluations (up to 6 months))

研究者

发起方
Magic Health Inc. (d.b.a. Nolla Health)
申办方类型
Industry
责任方
Sponsor

研究点 (1)

Loading locations...

相似试验

Evaluation of One-Shot Vision Differential Diagnosis... | 临床试验