Measurable Residual Disease as Surrogate Endpoint in CLL Trials: Methodological Concerns Raised Over Meta-Analysis Findings
核心洞察
A meta-analysis of 43 trials and 9,628 subjects found MRD has weak trial-level surrogacy for PFS (Spearman R = −0.33; R² = 0.06) despite strong individual-level association (HR 3.67, 95% CrI 3.34–4.03).
Critics raise concerns about dataset composition, including unclear trial counts, heterogeneous MRD assessment timing, and problematic inclusion of MRD-guided treatment arms.
The pooling of first-line fixed-duration trials with relapsed/refractory and continuous BTKi therapy studies may undermine the validity of the headline correlation.
A recent meta-analysis published in Leukemia by Wang and colleagues, examining whether measurable residual disease (MRD) can serve as a reliable surrogate for progression-free survival (PFS) in chronic lymphocytic leukemia (搜索) (CLL) clinical trials, has drawn critical methodological scrutiny from fellow researchers. The original analysis, encompassing 43 trials and 9,628 subjects, concluded that while detectable MRD is associated with a significantly higher individual-level risk of progression or death (HR 3.67, 95% CrI 3.34–4.03), trial-level surrogacy of MRD for PFS is weak overall (Spearman R = −0.33; R² = 0.06). The authors urged that MRD validity as a surrogate should be interpreted cautiously in clinical and regulatory decision-making.
Now, in a detailed correspondence also published in Leukemia, researchers have raised several critical questions regarding the methodology and interpretation underpinning these conclusions, arguing that the findings may not fully capture the context-dependent nature of MRD as a surrogate endpoint.
Dataset Transparency and Composition Under Question
A primary concern centers on the composition and transparency of the dataset. The headline figure of 43 trials in the abstract is not readily reconciled with later-described datapoints: the methods describe an individual-level analysis of 4,636 subjects from 28 trials and a trial-level analysis of 29 therapy comparisons from 18 trials, while the CONSORT-style flow diagram reports "43 Analyzed" decomposed into "32 for individual level analysis" and "18 for trial level analysis."
The supplementary trial list contains more than 40 entries, with the supplement noting that some articles originate from the same clinical trial but were included separately because they reported different types of data. Further complicating matters, at least one trial that clearly contributes a data point—the MRD-guided ibrutinib–venetoclax arm of the FLAIR trial—does not appear in the supplementary trial list. The correspondents have requested clarification of the actual number of included trials and comparisons, as well as how overlap between trials might have been adjusted in the independence-adjusted weighting scheme.
Heterogeneity in Pooled Studies Raises Concerns
The pooled individual-level hazard ratio of 3.67 is derived from trials encompassing markedly different clinical contexts: first-line trials within a fixed-duration setting such as CLL14 and GLOW, relapsed/refractory disease studies such as MURANO, and single-arm phase 2 trials without comparator arms. These trials differ substantially in their included populations, therapy settings, and study phases, and in some cases included experimental or even ineffective regimens—such as mitoxantrone—which does not reflect an established CLL regimen.
Moreover, the association between end-of-treatment MRD status and PFS is strongly dependent on observation time. The correspondents note that "virtually all randomized trials with adequate post-treatment follow-up after contemporary fixed-duration treatments have demonstrated an MRD-PFS correlation." Notable exceptions include the AMPLIFY study, which has limited post-treatment follow-up and carries competing survival risk from COVID-19-related PFS events, and the CLL17 study, in which landmark MRD analyses were not reported due to limited post-treatment follow-up at the time of the primary analysis.
Problematic Inclusion of MRD-Guided Treatment Arms
A particularly pointed concern involves the inclusion of the MRD-guided ibrutinib–venetoclax arm of the FLAIR trial in the trial-level analysis. By trial design, treatment duration in this arm was determined by MRD status itself. The correspondents argue that "including such a comparison in an analysis intended to validate MRD as a surrogate for PFS appears problematic," and have requested the authors' rationale for retaining this comparison, as well as an indication of whether the headline correlation is materially altered by its exclusion.
MRD Assessment Timing Heterogeneity
The headline correlation of R = −0.33 is derived from 29 comparisons in which the definition and timing of MRD assessment differ inherently across trials. Some report landmark PFS results from end-of-treatment MRD, such as CLL14 and CLL13, while others report on-treatment MRD during continuous therapy. The correspondents emphasize that "end-of-treatment MRD after fixed-duration venetoclax-based therapy and on-treatment MRD during continuous BTK (搜索)-inhibitor therapy reflect different points in fundamentally different disease trajectories," questioning whether pooling them into a single primary analysis is appropriate.
While supplementary sensitivity analyses do stratify by MRD assessment at 9 and 15 months from start of therapy, these remain inherently different between treatment paradigms and do not inform the primary trial-level estimate on which the main conclusion rests.
The IQWiG/Buyse Framework and BTKi-Specific Concerns
Finally, the correspondents cite the IQWiG/Buyse framework for establishing surrogate endpoint validity, which requires individual-level surrogacy to be established before trial-level surrogacy can be meaningfully assessed. Critically, individual-level surrogacy has not been established for continuous BTK (搜索)-inhibitor monotherapy—long-term follow-up of the E1912 trial demonstrated that detectable MRD does not preclude prolonged PFS in ibrutinib-treated patients, and this trial is itself among those included in the dataset.
Comparisons involving continuous BTKi therapy—including E1912, FLAIR IR versus FCR, HELIOS, GENUINE, and ELEVATE-TN—account for a substantial proportion of the trial-level data points despite this unresolved prerequisite. The correspondents warn that "reporting a weak correlation in a setting where the individual-level link itself has not been established risks misattributing the claim of missing surrogacy."
The debate underscores the complexity of validating surrogate endpoints in CLL, particularly as the treatment landscape evolves with both fixed-duration combination regimens and continuous BTK (搜索)-inhibitor therapies. The resolution of these methodological questions may have significant implications for future CLL trial design and regulatory decision-making.
