Synthetic Control Arms in Hematology: Generative AI Poised to Reshape Cancer Trial Design
核心洞察
Generative AI can create synthetic patient cohorts that reproduce the statistical and biological properties of real data, potentially replacing or complementing control arms in clinical trials.
Synthetic AML cohorts have retrospectively reproduced outcomes of the phase 2 SORAML trial and the GIMEMA AML1310 trial, matching original survival dynamics.
The EMA has approved synthetic and historical data for the control arm of the Phase II GD2-CART01 trial in pediatric neuroblastoma (搜索), setting a key regulatory precedent.
Developing new cancer drugs demands deep financial resources and extraordinary patience: the budget to bring a single cancer drug to market is estimated to exceed a billion US dollars, spent over nearly a decade, and the undertaking still frequently ends in failure. One of the top reasons cancer trials falter is recruitment failure, a problem that will only worsen as cancer entities fragment into molecularly-defined subgroups addressed by targeted agents in the era of precision oncology. When multiple trials compete for the same shrinking pool of eligible patients, trials begin to "cannibalize" each other, slowing accrual and increasing the likelihood of failure. This challenge raises a provocative question: if a standard-of-care control group remains a practical and regulatory necessity, can researchers create one instead of recruiting one?
Generative artificial intelligence (AI) can now produce synthetic data that are statistically indistinguishable from real data, spanning text, images, audio, video, and tabular data. The underlying principle is a statistical mimicry of feature distributions in real data, approximated by neural networks to create synthetic samples that follow the same distributions without being exact copies of their training data. This concept differs from digital twins, which constitute a digital replica of an individual patient to dynamically model trajectories given interventions. Synthetic cohort data, by contrast, reproduce feature distributions across a cohort, capturing statistical properties and inter-feature relationships rather than mirroring a single patient. Frequently used model architectures include generative adversarial networks (GANs), variational autoencoders, diffusion models, and transformers.
Synthetic Cohorts Reproduce Real Trial Outcomes in Hematology
In hematology, AI increasingly supports physicians in diagnosis and therapeutic decision-making. Researchers have demonstrated that synthetic data faithfully reproduce disease properties observed in real patient cohorts and may potentially substitute or replace control cohorts in clinical trials. In one example, GAN models were trained to generate synthetic bone marrow smears that experts failed to distinguish from real samples, and these synthetic images were used to train highly accurate image classifiers for leukemia detection in microscopy.
More directly relevant to trial design, researchers generated synthetic patients trained on 1,606 acute myeloid leukemia (搜索) (AML) patients from previous clinical trials of the German Study Alliance Leukemia, using multimodal tabular data including clinical, laboratory, and genetic features. The resulting synthetic cohorts recapitulated real patient properties, disease biology, and survival dynamics. These synthetic AML patients were then used to retrospectively replace the control cohort of the phase 2 trial SORAML, which evaluated the addition of sorafenib to standard induction therapy, effectively reproducing original trial outcomes when comparing the original intervention to the synthetic control arm. Similarly, Piciocchi et al. generated synthetic AML patients based on a cohort of 500 patients from the GIMEMA AML1310 trial, reporting survival dynamics matching original trial outcomes. In myelodysplastic neoplasms (搜索) (MDS), D'Amico et al. created a large synthetic cohort to investigate prognostically relevant features, which overlapped with the IPSS-M features, suggesting the IPSS-M could have also been discovered in a synthetically augmented cohort.
A Regulatory Precedent in Rare Disease
The feasibility of synthetic control arms has been historically confirmed by the European Medicines Agency (搜索) (EMA), which approved the use of synthetic and historical data to establish the control arm in the Phase II GD2-CART01 clinical trial, conducted by the Bambin Gesù Hospital (搜索) in Rome. That trial assesses the efficacy of anti-GD2 CAR-T cells (搜索) in pediatric patients with relapsed or progressive neuroblastoma (搜索), a rare and highly serious cancer. This "green light" sets a crucial precedent for the scientific community, demonstrating that virtual arms can effectively complement research where traditional sampling is impractical or ethically problematic.
The context of rare diseases underscores the urgency. Rare diseases are conditions with an incidence of fewer than 6 cases per 100,000 people per year, yet collectively account for 25 per cent of cancer diagnoses. Traditional randomized controlled trials (RCTs) face significant limitations in these settings: recruiting sufficient patients is often impossible due to small eligible populations, and assigning subjects with serious unmet needs to a control arm treated with ineffective therapies or a placebo raises profound ethical dilemmas. The under-representation of the elderly population and the rapid evolution of the standard of care further risk rendering control arms obsolete before a study is completed.
Building the Infrastructure for Synthetic Data
At the forefront of this transition is the IRCCS Istituto Clinico Humanitas (搜索), which has been developing synthetic data models since 2021, beginning with oncohaematology and subsequently expanding into neurology, gastroenterology, hepatology, and rheumatology. Through the Humanitas AI Centre, the first integrated AI research centre within an Italian IRCCS, clinicians, engineers, and data scientists collaborate to create robust, reproducible, and fully explainable algorithms. From this ecosystem, the spin-off Train was established in 2023, focusing on the development of digital twins and synthetic datasets that remain secure within the hospital's infrastructure. Data quality is continuously monitored by the SAFE (Synthetic vAlidation FramEwork) framework, developed as part of the European Horizon 2020 projects SYNTHIA and SYNTHEMA, which assesses statistical fidelity and clinical utility through transparent verification processes.
Pitfalls and the Path to Clinical Implementation
Despite their potential, synthetic patient data are no panacea. First, data generation is locked in a conundrum: to train a generative model, one needs a large, diverse, and representative training sample so the model can infer accurate feature distributions and capture intricate relationships. Generative models cannot make up useful data from scratch. For clinical trials, this may best apply to patients receiving standard of care, since plenty of patient records for training exist across healthcare systems, previous trials, and registries.
Second, generative models mirror feature distributions of their training data and may carry over or even amplify implicit biases, including local patient demographics, institutional preferences in treatment selection, or overrepresented subgroup properties. Training data should therefore come from multiple sources, ideally including real-world registries in addition to prior clinical trial data, to ensure adequate sample size, representativeness, and generalizability. The inclusion of minority groups is of particular importance, as otherwise generative models will not contribute to solving the equity problem in medicine but will quietly intensify it.
Third, privacy preservation is not guaranteed by default. Sensitive information can still be exposed through unintended model behavior or adversarial manipulation, such as membership inference or model inversion attacks. Safeguarding patient privacy should guide the data generation process from inception, accounting for a privacy-usability tradeoff: the more synthetic data differ from original data, the more privacy-compliant they become, yet the less useful they are for downstream tasks. Differential privacy budgets, real-to-synthetic distance metrics, and design safeguards combatting adversarial attacks help mitigate the risk of privacy breaches.
Lastly, synthetic data remain in a regulatory limbo. Regulatory agencies increasingly acknowledge the need for alternative control cohorts in settings where placebo control is not feasible or recruitment is slowed by small or inaccessible patient populations, especially in rare cancers or molecular subgroups. However, no regulatory framework currently exists for synthetically controlled trials. Regulators will have to define appropriate quality measures, including transparency requirements on training cohort properties, disclosure of model architecture, metrics for fidelity and usability, and privacy preservation. Critically, synthetic data generation allows a degree of customization that could lead to a conflict of interest, as entities with commercial or academic stakes in a trial outcome could potentially "cherry-pick" a synthetic control cohort to manufacture a desired result. Regulatory agencies should therefore provide a framework for synthetic data generation by independent third parties, with the resulting cohort withheld from investigators and sponsors until completion of intervention arm data collection.
In summary, synthetic data hold the potential to reduce barriers in data sharing and may enable novel trial designs, accelerate recruitment, reduce failure rates, and provide more patients with access to investigational therapies, particularly in the era of precision therapies in hematology. Yet they are no silver bullet: rigorous evaluation, quality assessment, privacy preservation, and regulatory guidance are needed before clinical implementation.
