Multimodal AI in Surgical Oncology Shows Promise but Lacks Prospective Evidence, Review Finds
核心洞察
A structured narrative review in Frontiers in Surgery assesses multimodal deep learning across preoperative, intraoperative and postoperative surgical oncology, concluding most evidence remains retrospective or single-center.
Radiogenomic models report external-validation AUCs up to 0.941 for noninvasive MSI (搜索) prediction in colorectal cancer (搜索), but retrospective design and site-specific imaging limit transportability.
The randomized RIDERS trial reduced positive margins at the neurovascular bundle in prostatectomy, yet attrition and short follow-up leave recurrence and functional benefit uncertain.
Multimodal deep learning (MDL) is advancing rapidly across the surgical oncology continuum, but fragmented clinical data, limited interpretability, unstable performance under distribution shift, and insufficient prospective evidence continue to constrain clinical adoption, according to a structured narrative review published in Frontiers in Surgery.
The review, authored by GuoBin Pan, synthesizes English-language literature published from January 2017 through January 2026 across PubMed/MEDLINE, Web of Science, and Scopus, prioritizing multicenter, external-validation, prospective, and interventional studies. It concludes that current evidence supports biologically informed AI as a promising decision-support layer, but that prospective multicenter trials, cost-effectiveness analyses, and governance safeguards are required before routine clinical use.
The Preoperative Gap: Anatomy Without Biology
The review frames the central problem in surgical oncology as a temporal gap: margin assessment remains predominantly morphological, while mutational profiles, proliferative activity, and tumor-microenvironment features are often unavailable until tissue sampling or postoperative analysis. This delay may contribute to missed occult disease or unnecessary removal of functional tissue.
Radiogenomic models attempt to close that gap by linking imaging phenotypes with molecular features. In locally advanced colorectal cancer (搜索), pretreatment dMMR (搜索)/MSI (搜索)-H status can identify candidates for PD-1 (搜索) blockade and potential organ-preservation strategies. Vision Transformer models have reported external-validation AUCs up to 0.941 for noninvasive MSI prediction, though the review notes that retrospective design, cohort enrichment, and site-specific imaging or histology may limit transportability.
Additional retrospective discrimination metrics cited include an AUC of 0.893 for dual-layer spectral CT iodine features associated with MSI (搜索)-H status, and an AUC of 0.885 for a model combining CT and whole-slide histology. The review cautions that these metrics do not establish calibration, treatment benefit, or safe selection for organ preservation.
For outcome prediction, an immunophenotype-guided radiomics model inferred immune-cell patterns with an AUC of 0.895 for predicting absence of pathological complete response, while a patho-radiomic model combining CT and histology achieved a concordance index of 0.82 for overall-survival prediction, exceeding TNM staging within the study cohort. Both findings are described as hypothesis-generating.
In occult metastasis detection, a DenseNet121 radiomics model combining metabolic and anatomical features discriminated malignant from benign mediastinal nodes in an external cohort in non-small cell lung cancer (搜索), supporting further evaluation as a triage aid but not justifying omission of guideline-recommended tissue confirmation.
Intraoperative Navigation: Feasibility, Not Superiority
Intraoperative applications reviewed include real-time augmented reality (AR), optical biopsy, and deformable image registration. The review emphasizes that real-time deployment is constrained by end-to-end latency, not inference alone, and that operating-room use requires synchronized video acquisition, local GPU or edge acceleration, secure high-bandwidth networking, low-latency display, and fail-safe fallback.
In a prospective usability study of 13 laparoscopic liver procedures, surgeons preferred less visually obstructive overlays, but the displays were not used to guide surgery. In a retrospective propensity-weighted cohort of 33 AR-assisted and 212 control procedures, conversion for failed tumor localization occurred in 0 versus 6 cases, but the difference was not statistically significant, operative time increased by approximately 10%, and resection margins and complications were unchanged.
The prospective randomized RIDERS trial enrolled 136 patients (133 treated), with only 82 reaching 12-month follow-up. AI-driven 3D AR guidance reduced positive margins at the preserved neurovascular bundle compared with cognitive MRI guidance, but the review states that attrition and short follow-up limit conclusions regarding biochemical recurrence and functional recovery. In neurosurgery, a retrospective study of 115 patients (39 mixed-reality and 76 standard-navigation procedures) reported smaller residual volumes but no overall-survival benefit, with nonrandom allocation limiting causal inference.
For rapid intraoperative diagnosis, RapidLymphoma was evaluated prospectively across four centers in 160 patients and in two independent cohorts (n = 420 and n = 59), achieving balanced accuracies of 97.81%, 95.44%, and 95.57% within 3 minutes. Only 25 patients in the prospective cohort had primary CNS lymphoma (搜索), and the review notes that performance at tertiary centers does not establish outcome benefit or transferability to lower-volume settings.
Optical margin assessment remains investigational. Hyperspectral segmentation models trained on 712 ex vivo datacubes achieved 0.98 pixel-level accuracy and 0.93 tumor recall, but the simulated approximately 10-minute workflow was not a prospective intraoperative outcome study. In a single-center prospective feasibility study of 22 prostatectomy patients (121 surgical-bed assessments), stimulated Raman histology interpretation achieved 98% accuracy, 83% sensitivity, and 99% specificity and detected 43% of patients with positive margins intraoperatively; the convolutional neural network was tested retrospectively in only 10 consecutive patients.
Postoperative Risk Stratification and Adjuvant Decisions
Postoperatively, pathomic and multi-omic frameworks may support biology-informed risk stratification. In colorectal liver metastases (搜索), an AI-assisted classifier reported AUC above 0.98 for histological growth patterns and increased pathologist accuracy from 75% to 92% in a human-in-the-loop setting, though multicenter reproducibility and decision impact remain to be established.
The review outlines a clinically governed human-in-the-loop workflow in which the surgeon or pathologist overrides the recommendation when input quality is inadequate, confidence is below a validated threshold, key modalities are missing, or the output conflicts with anatomy, guidelines, or patient goals. Overrides and outcomes should be logged for audit, and system failure should default to standard care.
On adjuvant therapy, the review notes that the average survival benefit of adjuvant chemotherapy in stage II colorectal cancer (搜索) is modest and heterogeneous, and that retrospective AI models have identified subgroups with different prognoses or treatment associations but do not establish who can safely omit chemotherapy. In hormone receptor-positive breast cancer (搜索), Orpheus (搜索) predicted a high Oncotype DX recurrence score with an AUC of 0.89 versus 0.73 for a clinicopathological nomogram and showed prognostic value among patients with scores of 25 or lower; RSClinN+ (搜索) refined risk estimates in node-positive disease. Both require prospective utility and cost-effectiveness studies before replacing genomic assays.
Large Language Models: Concordance Without Outcomes
The review also examines large language models (LLMs) in perioperative decision support. Retrospective case-based evaluations reported high concordance between general-purpose LLMs and breast-cancer tumor boards for guideline-concordant scenarios, but performance was lower for surgical judgment, genomic-test indications, and individualized risk-benefit trade-offs. In complex pancreatic-cancer scenarios, GPT-4 (搜索) achieved 61.5% accuracy for neoadjuvant-treatment planning. A meta-analysis of 56 studies across 15 cancer types reported overall accuracy of 76.2% and diagnostic accuracy of 67.4%, while most studies did not evaluate patient-safety risks.
For documentation, a HIPAA-compliant framework extracting structured variables from 45 rectal-cancer operative reports achieved performance comparable to a surgical resident and reduced processing time from more than two hours to approximately 10 minutes. A retrospective breast-cancer study found GPT-4 (搜索) extracted 26 registry variables with 97.1% overall accuracy and reduced median processing time from 15.5 to 1.2 minutes. A multicenter study of GPT-4o (搜索) for Chinese- and English-language lobectomy records reported accuracy of 0.966 and an F1 score of 0.882.
Vision-language models for intraoperative guidance remain predominantly experimental. The review notes that specialized architectures such as Surgical-MambaLLM and Surgery-R1 lack prospective operating-room validation, and that outputs may be unreliable under occlusion, smoke, bleeding, unfamiliar anatomy, or distribution shift.
From Discrimination to Utility
The review's central methodological argument is that technical discrimination should not be equated with clinical utility. Most reviewed studies are retrospective or single-center and report internal AUC or accuracy rather than calibration, treatment change, safety, or patient outcomes. Clinical readiness, the author writes, requires independent external validation, prospective silent testing, and then interventional evaluation with predefined failure modes, subgroup analyses, clinician-AI interaction, and workflow or patient outcomes.
Barriers to translation include biased training cohorts, weak external validation, automation-related patient harm, privacy or cybersecurity breaches, opaque reasoning, and uncertain responsibility when errors occur. The review also highlights unequal access, noting that genomic sequencing, whole-slide scanners, stimulated Raman histology or AR systems, and on-premises accelerators require capital, maintenance, annotation expertise, and cybersecurity support. Federated and swarm learning may reduce data silos without centralizing raw data, but do not remove selection bias, protocol heterogeneity, or the need for external validation.
Because errors may arise from model design, local integration, data shift, or clinician use, responsibility may be distributed among developers, institutions, and clinicians. The review calls for versioned audit logs, defined oversight, incident review, and local medico-legal assessment, concluding that AI output should remain advisory rather than replace accountable clinical judgment.
