JMML has never been a disease where the mutation alone can be trusted to tell the whole story. Two children may carry familiar RAS-pathway lesions and still travel very different clinical paths. One may settle into a more indolent course; another may relapse after maximal therapy. Even hematopoietic stem cell transplant, the central curative option for many patients, does not remove the uncertainty: approximately 50% of children relapse after transplant.
That is the clinical pressure behind this line of research. The useful question is not whether a model can sound sophisticated. It is whether a classifier can separate children into biologically meaningful risk groups well enough to change treatment intensity without pretending that retrospective performance is the same thing as bedside certainty.

DNA methylation classification mattered because it gave the field something genetics had not reliably provided: a reproducible way to organize prognosis across low, intermediate, and high-risk biology. The point was not that methylation replaced clinical judgment. The point was that methylation captured a layer of disease behavior that conventional mutation categories often missed.
Why methylation became the organizing signal
The first convincing step came from discovery work showing that JMML could be divided into methylation-defined groups with sharply different outcomes. Stieglitz and colleagues studied 79 patients using Illumina 450k methylation arrays and identified low, intermediate, and high methylation clusters. The 4-year event rates were 6% for low methylation, 45% for intermediate methylation, and 61% for high methylation. The same study also identified a methylation signature associated with spontaneous resolution, an especially important observation in a disease where overtreatment and undertreatment both carry real consequences.[1]
Lipka and colleagues, also in 2017, added a second piece of groundwork. In a 129-patient analysis, they showed that RAS-pathway mutation patterns were associated with epigenetic subclasses, helping connect genotype with methylation state rather than treating the two as unrelated layers of information.[2] These studies did not settle how methylation should be operationalized in routine risk assignment, but they made it hard to ignore methylation as a prognostic axis.
The uncomfortable lesson was that mutation-informed prognosis and methylation-informed prognosis were not interchangeable. For clinicians deciding whether to intensify therapy, that distinction is not academic. A risk classifier has to be more than associated with outcome; it has to be portable, interpretable enough to audit, and stable enough to survive a change in platform.
The consensus XGBoost classifier: what it actually classifies
The major consolidation came in 2021, when Schönung and colleagues built an international consensus classifier for JMML methylation subgroups. This is the study that turned scattered methylation signals into a more practical three-tier system. The model used XGBoost, a gradient-boosted decision tree method, to assign patients to low, intermediate, or high methylation groups using 124 CpG sites selected from the 5,000 most variable CpGs.[3]
The scale matters. The classifier was developed from pooled international data including 292 patients, then tested in an additional 47-patient validation cohort. It predicted methylation subgroups with 98% accuracy across EPIC methylation arrays and targeted MethylSeq platforms.[3] That number is impressive, but it should be read carefully: it is retrospective pooled validation, not proof that every center can deploy the tool prospectively tomorrow with identical performance.
| Feature | Consensus classifier |
|---|---|
| Model type | XGBoost methylation subgroup classifier |
| Input features | 124 CpG sites selected from the 5,000 most variable CpGs |
| Risk groups assigned | Low methylation, intermediate methylation, high methylation |
| Development and validation material | 292 pooled patients plus 47 validation patients |
| Reported accuracy | 98% across EPIC arrays and targeted MethylSeq |
| Clinical importance | Methylation subgroup retained independent prognostic significance for overall survival |
The cross-platform point is not a technical footnote. Genome-wide arrays are useful for discovery and deep characterization, but they are not the same as a focused clinical workflow. The Schönung classifier was evaluated across EPIC arrays and targeted MethylSeq, which is exactly the kind of translation step that matters if a methylation model is expected to inform treatment assignment rather than sit permanently inside retrospective datasets.[3]
The more important result was not the headline accuracy. In multivariable analysis, methylation subgroup was the only factor that retained independent prognostic significance for overall survival, with p=0.046 for the XGBoost classifier and p=0.039 for the MethylSeq classifier.[3] That is the reason the model deserves attention from trial designers. It did not merely reproduce a known clinical label; it carried prognostic information when other variables were considered.
This is also where the language around AI needs discipline. The classifier does not make JMML simple. It does not erase transplant risk, erase therapeutic uncertainty, or prove that treatment escalation based on methylation will improve survival. It gives a validated molecular framework for stratification, and that is already a large advance in a disease where risk has too often been inferred from incomplete proxies.

Why independent prognostic value changes the conversation
A model that agrees with prior labels may be neat. A model that independently predicts survival is clinically disruptive. In JMML, that distinction matters because treatment decisions are not harmless sorting exercises. Low-risk assignment may support a less intensive approach. High-risk assignment may justify adding more therapy before or around transplant. Intermediate-risk assignment may be the most uncomfortable category, because it prevents the false reassurance of a binary system.
The consensus classifier’s three-tier structure is therefore more clinically honest than a simple high-versus-low split. It leaves room for patients whose biology does not declare itself cleanly at the extremes. That is not a weakness of the system; it is a better reflection of the disease.
For a molecular pathologist, the 124-CpG design is also reassuring for a practical reason. It narrows a genome-wide signal into a defined feature set that can be interrogated and transferred. A black-box label with no durable molecular footprint would be much less useful. The field does not need another beautiful retrospective classifier that falls apart when the assay changes.
Clinical-parameter models are useful bridges, not substitutes
There is an obvious reason to ask whether methylation subgroup can be approximated from ordinary clinical data. Not every center can obtain methylation profiling quickly, and not every health system can absorb specialized assays on demand. Imaizumi and colleagues tested machine learning models using only clinical parameters: age, fetal hemoglobin, platelet count, and PTPN11 status. In 128 training patients and 72 validation patients, decision tree, support vector machine, and naïve Bayes approaches reached 82% to 84% concordance with array-based methylation classification.[4]
That is attractive if the alternative is no methylation-informed estimate at all. It is not attractive enough to confuse with a methylation classifier. Concordance of 82% to 84% means roughly 16% to 18% of patients would be assigned differently than with array-based methylation classification.[4] In risk-adapted leukemia care, that is not a small clerical inconvenience. It can mean a child is treated as biologically lower or higher risk than the methylation reference would suggest.
| Approach | Practical appeal | Main limitation |
|---|---|---|
| Genome-wide methylation arrays | Rich molecular profiling and discovery-grade context | Less convenient for rapid routine decision-making |
| 124-CpG XGBoost consensus classifier | Validated three-tier methylation assignment across EPIC arrays and targeted MethylSeq | Retrospective validation still needs prospective multi-center confirmation |
| Clinical-parameter ML models | Accessible where methylation profiling is unavailable | Only 82% to 84% concordant with array-based methylation classification |
| BMP4 single-locus assay | Fast, simplified methylation readout | A single-locus biomarker is a trade-off against full consensus classification |
The clinical-parameter models are best understood as bridges. They may help centers think probabilistically when methylation testing is unavailable, delayed, or logistically impossible. They should not be presented as equivalent substitutes for a methylation classifier, because the evidence does not support that stronger claim.
The BMP4 assay sits in the middle
A different compromise is the single-locus methylation assay. Ghanjati and colleagues evaluated BMP4 methylation by bisulfite next-generation sequencing in 111 patients, reporting a 48-hour turnaround and 0.89 specificity compared with array-based methylation classification. Five-year disease-free survival was 38% in the high-methylation group and 62% in the normal-methylation group, with p=0.007.[5]
The attraction is obvious: a faster and simpler assay fits the tempo of clinical decision-making better than a workflow that may take weeks. But the trade-off is just as obvious. BMP4 is a simplified biomarker, not the full 124-CpG consensus classifier. It may be useful when speed and feasibility dominate, but it should not be inflated into the same evidentiary category as the cross-platform XGBoost system.
From stratification to treatment assignment: the TRAZA trial
The decisive test for this field is not whether methylation subgroups can be named. It is whether assigning therapy by those subgroups improves what happens to children. That is why the St. Jude TRAZA trial is the natural endpoint of the current evidence chain. Active in 2026, TRAZA is the first prospective trial to assign JMML therapy based on methylation subgroup risk stratification.[6]
In TRAZA, low-risk patients receive azacitidine plus trametinib, while high-risk patients receive the same combination plus fludarabine and cytarabine.[6] That design moves methylation classification from prognostic description into treatment allocation. It is the point at which a classifier stops being merely a way to explain past outcomes and becomes part of a prospective therapeutic strategy.
This is the right level of ambition. The consensus methylation classifier is currently the strongest prognostic framework in JMML and a serious candidate for guiding risk-adapted therapy. But prospective use is the standard that matters now. Retrospective 98% accuracy is not the same as prospective multi-center reliability, and clinical trial stratification is not the same as a cleared standalone treatment device.
For now, the defensible position is cautious: DNA methylation machine learning has made JMML risk stratification more biologically legible and more clinically useful. The field is now testing whether that framework can safely guide treatment intensity before it is called settled practice.
References
- The genomic landscape of juvenile myelomonocytic leukemia, Nature Communications, 2017.
- RAS-pathway mutation patterns define epigenetic subclasses in juvenile myelomonocytic leukemia, Nature Communications, 2017.
- International consensus definition of DNA methylation subgroups in juvenile myelomonocytic leukemia, Clinical Cancer Research, 2021.
- Machine learning approach for predicting DNA methylation classification of juvenile myelomonocytic leukemia using clinical parameters, Scientific Reports, 2022.
- BMP4 methylation as a simplified biomarker for juvenile myelomonocytic leukemia risk stratification, Clinical Epigenetics, 2025.
- TRAZA: Trametinib and Azacitidine for Juvenile Myelomonocytic Leukemia, St. Jude Children’s Research Hospital.
Comments
Join the discussion with an anonymous comment.