Review Maps Path From Foundation Models to Accountable Clinical AI in Medical Imaging
核心洞察
A July 2026 review from Chinese Academy of Sciences and Beihang University (搜索) researchers outlines four development paths for medical imaging foundation models (搜索), from image-representation pre-training to dynamic visual sequence modelling.
The review argues the field's next milestone should be evidence that a model works safely within a defined clinical role, not further increases in parameter counts.
A separate structured critical review of 80 sources concludes LLM-based clinical decision support is not supported for routine autonomous use and requires mandatory clinician supervision.
Medical imaging foundation models (搜索) are advancing rapidly in capability, but the evidence needed to place them safely inside clinical workflows remains thin, according to two 2026 reviews that converge on a common conclusion: the binding constraint on clinical impact is no longer model performance but validation, governance and human-AI interaction.
Researchers from the Institute of Automation, Chinese Academy of Sciences, and the School of Engineering Medicine at Beihang University (搜索) published their review in July 2026 in the Medical Journal of Peking Union Medical College Hospital (搜索) (DOI: 10.12290/xhyxzz.2026-0416). The authors examine how medical imaging foundation models (搜索) are built, adapted, evaluated and moved toward clinical use, assessing progress across radiology, digital pathology, ultrasound and surgical video. They frame the central question not as whether these models can perform many tasks, but whether they can deliver stable, verifiable value in real-world clinical environments.
Four Development Paths, and the Limits of Data Volume
The review organizes the field into four overlapping development paths: image-representation pre-training, image-language alignment, integration of multiple clinical data sources, and modelling of dynamic visual sequences. Training data may include computed tomography (CT), magnetic resonance imaging (MRI), X-rays, ultrasound, pathology slides, reports, laboratory measurements, treatment records, endoscopy and surgical video.
The authors caution that data volume alone can be misleading, because millions of image patches or video frames may not correspond to an equivalent number of independent patients. Quality control, deduplication, patient-level independence, cross-centre coverage and reliable pairing between images and text are therefore described as critical considerations.
Adaptation strategies range from lightweight task heads to fine-tuning, prompt learning and instruction tuning. Single-modality examples span cancer subtyping, mutation prediction, survival estimation and lesion segmentation, while vision-language systems add retrieval and question answering.
Evaluation Beyond Accuracy
The review proposes that evaluation extend beyond accuracy to three additional dimensions: algorithmic robustness under external data and input disturbances; clinical usefulness assessed through comparisons of clinician-only versus clinician-model performance; and workflow outcomes including reporting time, triage efficiency, repeat examinations, resource use and patient outcomes. It also stresses that a foundation model may still underperform a task-specific system in a clearly defined clinical setting.
The authors argue the field's next milestone should not be another increase in parameter counts, but evidence that a model can work safely within a defined clinical role. A more feasible path, they suggest, may be to use foundation models to assist with limited, reviewable tasks — triage, report drafting, interactive segmentation, risk stratification or structured follow-up — rather than attempting to replace an entire diagnostic process.
For deployment, the review outlines requirements including clear indications, prohibited uses, input-quality requirements, uncertainty signals, human review responsibilities, failure reporting, version tracking and revalidation after updates. Models should connect reliably with picture archiving and communication systems (PACS), radiology information systems (RIS) and hospital information systems (HIS), while preserving logs of inputs, outputs, clinician edits, warnings and software versions. Governance must address privacy, consent, secondary data use, copyright, demographic bias, performance drift and responsibility for errors. The wider implication, the authors write, is that clinical translation will depend on a layered partnership among general foundation models, specialty-specific systems and human oversight, with each component serving a clearly bounded and auditable role.
A Parallel Appraisal of Deep Learning and Generative AI
A structured critical narrative review published in Frontiers in Digital Health reaches similar conclusions from a different evidence base. Drawing on 80 sources — 37 screened studies plus 43 landmark primary studies, architectural papers and clinical-AI reporting standards published between 2014 and 2026 — the authors trace the evolution from convolutional neural networks (搜索) (CNNs) to U-Net (搜索) and encoder-decoder networks, then to Vision Transformers (搜索) (ViTs), generative adversarial networks (搜索) (GANs), diffusion models (搜索), large language models (搜索) (LLMs) and retrieval-augmented generation (搜索) (RAG).
Each model family is compared across nine dimensions: input modality, task, data requirements, validation level, interpretability, failure modes, clinical readiness, regulatory considerations and human-oversight need. Supervised deep learning reaches clinically useful performance across CT, MRI and pathology on well-scoped detection, segmentation and classification tasks, though the authors note results are task-, dataset- and site-dependent and prospective evidence is limited. Reported indicative performance includes AUROC of approximately 0.85–0.97 for CNN classification on specific radiology tasks and Dice of approximately 0.80–0.95 for organ segmentation with encoder-decoder architectures.
Generative models can produce synthetic images or cross-modality translations to address data scarcity, but the review warns they may amplify hidden dataset biases and generate anatomically incorrect images. Diffusion models (搜索) are described as offering superior mode coverage versus GANs but at high inference cost and with limited clinical validation. LLM-based clinical decision support systems show promise for guideline-concordant reasoning and medication-safety checks, yet still face hallucination, calibration and regulatory uncertainty.
Evidence Does Not Support Autonomous LLM Decision Support
The review states plainly that, on current evidence, LLM-based clinical decision support systems are not supported for routine autonomous use and require clinician supervision. It cites several lines of evidence. In a synthesis of clinical-LLM literature screening 4,609 papers from 2022 to 2025, few studies used real-world patient data and only 19 were prospective, randomised, controlled trials.
In one evaluation, state-of-the-art LLMs tested against 2,400 real intensive-care-unit patient cases from four abdominal diseases outperformed physicians in accuracy but failed to adhere to medical guidelines, failed to interpret laboratory results and were sensitive to the order and amount of information presented; the authors concluded that current LLMs are not ready for autonomous clinical decision making.
Head-to-head evidence favored augmentation over substitution. In a prospective, cross-over study of 91 medication safety scenarios across 16 specialties, a pharmacist-plus-LLM combination achieved a 32% increase in drug-related problem detection compared with pharmacists alone, while the LLM alone performed less effectively than pharmacists. In gastroenterology and hepatology, LLM recommendations were consistent with guidelines during structured endoscopy and histology input but were limited to the level of trainees and below subspecialty hepatologists.
A meta-analysis of predictive AI-CDSS across 50 studies and 17 specialties showed moderate pooled discrimination overall, with relatively high specificity but lower and less consistent sensitivity, and most evaluations were performed on historical, retrospective data rather than tested in real clinical workflows. An audit of FDA-cleared AI devices found most were approved based on retrospective studies, with fewer than one in five providing multi-site or prospective evaluation.
Capability and Reliability Move in Opposite Directions
The review identifies a consistent pattern across the model continuum: as capability increases — reasoning over more varied data types and producing more flexible recommendations — reliability, interpretability and controllability decrease. The authors attribute this in part to the nature of probabilistic models that sample from learned distributions rather than following deterministic logic, describing it as a structural property rather than a temporary engineering problem.
The practical implication offered is task-dependent selection. For situations requiring action before a decision, such as sepsis management, stroke triage or ICU dosing, validated predictive deep-learning CDSS may be more suitable than flexible LLMs. LLM reasoning may be appropriate for lower-stakes tasks such as generating a first differential list, patient information or documentation, where the clinician reviews the output.
The authors also caution that traditional machine learning, deep learning and generative AI represent an expansion of the clinical AI toolkit rather than a linear progression in which each supersedes the last. They suggest the most effective deployed solutions are likely to combine all three, with each component confined to its validated operating range and subject to human oversight.
Reporting Standards as a Staged Evidence Pathway
Both reviews emphasize reporting and evaluation frameworks as the basis for credible clinical AI claims. The Frontiers review maps technologies to CONSORT-AI and SPIRIT-AI for randomised trials and protocols of AI interventions, TRIPOD+AI for prediction-model studies, CLAIM for AI in medical imaging, DECIDE-AI for early-stage live clinical evaluation, STARD-AI for diagnostic-accuracy studies, PROBAST+AI for risk-of-bias appraisal, and FUTURE-AI for trustworthiness principles across the AI lifecycle.
These standards suggest a staged evidence pathway: CLAIM- or TRIPOD+AI-compliant development reporting, then DECIDE-AI early clinical evaluation, then CONSORT-AI and SPIRIT-AI trials. The authors propose a human-in-the-loop architecture in which multimodal inputs are encoded and fused, grounded against a retrieval store, and passed through safety guardrails covering uncertainty and calibration, source attribution and out-of-scope abstention before mandatory clinician verification, with every step logged for audit and post-deployment monitoring.
The review sets out a staged roadmap. In the near term, it calls for reporting all new imaging and CDSS models according to CLAIM and TRIPOD+AI, publishing external-validation results and calibration measures alongside discrimination metrics, and deploying LLM-based CDSS only within a human-in-the-loop safety architecture. Over two to five years, it recommends DECIDE-AI-compliant early clinical evaluations followed by CONSORT-AI and SPIRIT-AI randomised trials evaluating clinician-AI team performance and patient outcomes, alongside federated multi-centre validation infrastructures and routine fairness auditing. Over five years and beyond, it calls for regulatory science for adaptive and generative systems, including non-inferiority margins and post-deployment drift monitoring, and for efficient, edge-deployable models to extend validated AI to community healthcare and resource-limited settings.
Regulatory Frameworks Lag Adaptive Systems
The review notes that the FDA's Software-as-a-Medical-Device guidance and the EU AI Act are largely based on the assumption of a "locked" algorithm, a model that does not fit continuously updated LLMs and foundation models whose performance can change between iterations. It also points to unresolved liability questions when a clinician implements or dismisses an AI recommendation that fails, and to documented risks of both under-reliance, where a clinician dismisses correct advice, and automation bias, where incorrect advice is accepted without questioning.
Additional barriers cited include the cost of real-time LLM-CDSS inference for resource-constrained providers, the need for local validation even of cleared tools, and the absence of resources for post-deployment model drift monitoring. The authors note that alarm fatigue from poorly set alert thresholds is a documented source of the gap between prospective and retrospective performance.
Both reviews acknowledge limitations. The Frontiers article is a structured critical narrative review, not a PRISMA-style systematic review; its 37-study core was selected to illustrate the field's development rather than to be exhaustive, it was not registered, and individual studies were not assessed with a formal risk-of-bias instrument, though a descriptive appraisal is provided. The imaging foundation model review similarly calls for deeper research into representative data, prospective validation, workflow integration and continuous governance.
The shared conclusion is that safe deployment, not raw capability, now limits clinical impact. As the Frontiers authors put it, model capability will continue to be supplied by commercial development, but validation, governance and human-AI interaction protocols will not.
