AI-Generated Synthetic Real-World Data Shows Promise for Accelerating Cancer Clinical Trials
核心洞察
Synthetic real-world data generated by AI can replicate survival outcomes and clinical patterns from real patient cohorts while protecting patient privacy, according to research presented at ESMO Congress 2025.
A classification and regression tree (CART) model demonstrated the highest fidelity to the source cohort of 19,164 metastatic breast cancer (搜索) patients, with re-identification risk below 2%.
Synthetic datasets can facilitate data sharing, optimize clinical trial design through external control arms, and enable patient matching for outcome prediction.
Researchers at Dana-Farber Cancer Institute have demonstrated that AI-generated synthetic real-world data can faithfully reproduce survival outcomes and multivariate clinical patterns from large patient cohorts, offering a potential pathway to accelerate clinical trials and drug development while safeguarding patient privacy. The findings were presented by Eddy Saad, MD, MSc, Research Fellow in Medicine at Dana-Farber Cancer Institute, during the European Society for Medical Oncology (ESMO) Congress.
The work addresses what Dr. Saad described as a "vicious cycle" created by the accelerating pace of cancer drug approvals by the U.S. Food and Drug Administration over the past decade. "With each new therapy, we uncover new unmet needs and these create an exponential demand for data…from clinical trials, and these require enormous resources, perhaps most importantly, time," he said.
Building Synthetic Cohorts from Real Patient Data
Synthetic real-world data are AI-generated datasets that replicate the statistical patterns and properties of real-world patient data without containing any individual patient identifiers. Dr. Saad and colleagues derived all AI-based synthetic cohorts from a source cohort of 19,164 patients with metastatic breast cancer (搜索) in the Flatiron Health (搜索) database who received first-line treatment between 2011 and 2023.
The team evaluated two distinct modeling approaches. The first was a conditional tabular generative adversarial network (CTGAN), which uses a dynamic interaction between a generator and a discriminator that train each other to produce synthetic data indistinguishable from the source cohort. To maintain model stability and enforce privacy constraints, the researchers applied noise weight clipping, generating five versions: CTGAN base, CTGAN ln, CTGAN low, CTGAN medium, and CTGAN high.
The second model was a classification and regression tree (CART), which learns relationships within the data to replicate clinical patterns and decision-making. "CTGANs can be a sort of black box, so we wanted to offer a more transparent, clinician-friendly approach," Dr. Saad explained. "These rely on a series of branching, if-then decisions, [similar to] how a clinician would reason through a case."
Validating Fidelity and Privacy
The six resulting models were evaluated for fidelity to the original dataset and for protection of patient privacy. Absolute standardized mean difference analysis showed that models with higher levels of privacy, including CTGAN medium and CTGAN high, deviated more from the original dataset, whereas the CART model showed minimal differences from the source cohort.
Kaplan-Meier curves for progression-free and overall survival outcomes demonstrated greater divergence from the source cohort among models with higher privacy levels, while the CART model closely overlapped with the source cohort. Multivariate analysis using regression models further confirmed that the AI-based models—particularly CART—recapitulated the original variables and correlations.
Regarding privacy, with an acceptable re-identification risk threshold of 9%, all models demonstrated risk levels of 2% or lower, with the CART model exhibiting the highest risk among the six.
"Synthetic real-world data can strike the optimal balance between data utility and privacy. They faithfully reproduce univariate survival and multivariate patterns and have an acceptable re-identification risk," Dr. Saad said. "Synthetic datasets can therefore be leveraged to facilitate data sharing, to optimize clinical trial design, and to improve clinical decision-making."
Clinical Applications and a Hypothetical Case
Dr. Saad outlined several potential clinical applications, including facilitating data sharing across multiple stakeholders globally, modeling real-world control arms for clinical trial design, and enabling patient matching with outcome prediction. With modern synthetic control arms, "we could compare a new regimen and perhaps even accelerate the design of drugs and approval of these drugs," he noted.
He presented a hypothetical case of a 63-year-old woman with HR-positive, HER2 (搜索)-positive metastatic breast cancer (搜索), an ECOG performance status of 1, and a body mass index of 27.4 kg/m². Her case was compared with 200 "nearest neighbors," who generally received one of four treatment options: an aromatase inhibitor plus a CDK4/6 (搜索) inhibitor (real-world progression-free survival [PFS] = 11.6 months); an aromatase inhibitor plus a HER2-targeted agent (PFS = 10.9 months); a HER2-targeted agent plus a taxane (PFS = 17.2 months); or an aromatase inhibitor, HER2-targeted agent, and a taxane (PFS = 24.2 months).
"Looking ahead, our vision is that of an integrated platform where we could continuously ingest real-world data from multiple sources, process them through these AI models, obtain their twin datasets, and use these to improve clinical research and clinical outcomes for patients," Dr. Saad concluded.
Expert Perspective: Promise and Caution
Julien Vibert, MD, PhD, of the Drug Development Department at Gustave Roussy (搜索) in Paris and an invited discussant at ESMO, provided additional context. He noted that synthetic real-world data differ from de-identified patient data: synthetic cohorts are generated by AI algorithms that restructure source data so that individual patients cannot be re-identified.
Dr. Vibert highlighted a fundamental tradeoff between fidelity to source data and preservation of patient privacy, with improvements in one often coming at the expense of the other. He noted that the Flatiron Health (搜索) database cohort—comprising more than 19,000 patients with metastatic breast cancer (搜索)—represents one of the largest synthetic datasets in oncology to date.
Although synthetic models can enable data sharing and facilitate external control arms, Dr. Vibert cautioned that synthetic datasets may amplify biases in the source data and carry risks of misuse or overconfidence if not properly validated. He added that synthetic real-world data can also pose risks related to security breaches and ethical concerns around data governance. These challenges have complicated the regulatory acceptance of synthetic cohorts to date.
"AI can potentially help us to design smarter trials," Dr. Vibert said, while emphasizing that AI-based synthetic cohorts cannot replace real-world data or clinical trials. Clinical validation, he stressed, will remain indispensable.
Dr. Saad reported receiving research funding from Genentech, EMD Serono, and OncoHost (搜索). Dr. Vibert reported no conflicts of interest.
