JMML has never been a disease where the mutation alone can be trusted to tell the whole story. Two children may carry familiar RAS-pathway lesions and still travel very different clinical paths. One may settle into a more indolent course; another may relapse after maximal therapy. Even hematopoietic stem cell transplant, the central curative option for many patients, does not remove the uncertainty: approximately 50% of children relapse after transplant.

That is the clinical pressure behind this line of research. The useful question is not whether a model can sound sophisticated. It is whether a classifier can separate children into biologically meaningful risk groups well enough to change treatment intensity without pretending that retrospective performance is the same thing as bedside certainty.

DNA methylation data flowing through a machine learning network into low, intermediate, and high JMML risk groups

DNA methylation classification mattered because it gave the field something genetics had not reliably provided: a reproducible way to organize prognosis across low, intermediate, and high-risk biology. The point was not that methylation replaced clinical judgment. The point was that methylation captured a layer of disease behavior that conventional mutation categories often missed.

Why methylation became the organizing signal

The first convincing step came from discovery work showing that JMML could be divided into methylation-defined groups with sharply different outcomes. Stieglitz and colleagues studied 79 patients using Illumina 450k methylation arrays and identified low, intermediate, and high methylation clusters. The 4-year event rates were 6% for low methylation, 45% for intermediate methylation, and 61% for high methylation. The same study also identified a methylation signature associated with spontaneous resolution, an especially important observation in a disease where overtreatment and undertreatment both carry real consequences.[1]

Lipka and colleagues, also in 2017, added a second piece of groundwork. In a 129-patient analysis, they showed that RAS-pathway mutation patterns were associated with epigenetic subclasses, helping connect genotype with methylation state rather than treating the two as unrelated layers of information.[2] These studies did not settle how methylation should be operationalized in routine risk assignment, but they made it hard to ignore methylation as a prognostic axis.

The uncomfortable lesson was that mutation-informed prognosis and methylation-informed prognosis were not interchangeable. For clinicians deciding whether to intensify therapy, that distinction is not academic. A risk classifier has to be more than associated with outcome; it has to be portable, interpretable enough to audit, and stable enough to survive a change in platform.

The consensus XGBoost classifier: what it actually classifies

The major consolidation came in 2021, when Schönung and colleagues built an international consensus classifier for JMML methylation subgroups. This is the study that turned scattered methylation signals into a more practical three-tier system. The model used XGBoost, a gradient-boosted decision tree method, to assign patients to low, intermediate, or high methylation groups using 124 CpG sites selected from the 5,000 most variable CpGs.[3]

The scale matters. The classifier was developed from pooled international data including 292 patients, then tested in an additional 47-patient validation cohort. It predicted methylation subgroups with 98% accuracy across EPIC methylation arrays and targeted MethylSeq platforms.[3] That number is impressive, but it should be read carefully: it is retrospective pooled validation, not proof that every center can deploy the tool prospectively tomorrow with identical performance.

FeatureConsensus classifier
Model typeXGBoost methylation subgroup classifier
Input features124 CpG sites selected from the 5,000 most variable CpGs
Risk groups assignedLow methylation, intermediate methylation, high methylation
Development and validation material292 pooled patients plus 47 validation patients
Reported accuracy98% across EPIC arrays and targeted MethylSeq
Clinical importanceMethylation subgroup retained independent prognostic significance for overall survival

The cross-platform point is not a technical footnote. Genome-wide arrays are useful for discovery and deep characterization, but they are not the same as a focused clinical workflow. The Schönung classifier was evaluated across EPIC arrays and targeted MethylSeq, which is exactly the kind of translation step that matters if a methylation model is expected to inform treatment assignment rather than sit permanently inside retrospective datasets.[3]

The more important result was not the headline accuracy. In multivariable analysis, methylation subgroup was the only factor that retained independent prognostic significance for overall survival, with p=0.046 for the XGBoost classifier and p=0.039 for the MethylSeq classifier.[3] That is the reason the model deserves attention from trial designers. It did not merely reproduce a known clinical label; it carried prognostic information when other variables were considered.

This is also where the language around AI needs discipline. The classifier does not make JMML simple. It does not erase transplant risk, erase therapeutic uncertainty, or prove that treatment escalation based on methylation will improve survival. It gives a validated molecular framework for stratification, and that is already a large advance in a disease where risk has too often been inferred from incomplete proxies.

Three-tier JMML methylation risk stratification with low, intermediate, and high methylation zones

Why independent prognostic value changes the conversation

A model that agrees with prior labels may be neat. A model that independently predicts survival is clinically disruptive. In JMML, that distinction matters because treatment decisions are not harmless sorting exercises. Low-risk assignment may support a less intensive approach. High-risk assignment may justify adding more therapy before or around transplant. Intermediate-risk assignment may be the most uncomfortable category, because it prevents the false reassurance of a binary system.

The consensus classifier’s three-tier structure is therefore more clinically honest than a simple high-versus-low split. It leaves room for patients whose biology does not declare itself cleanly at the extremes. That is not a weakness of the system; it is a better reflection of the disease.

For a molecular pathologist, the 124-CpG design is also reassuring for a practical reason. It narrows a genome-wide signal into a defined feature set that can be interrogated and transferred. A black-box label with no durable molecular footprint would be much less useful. The field does not need another beautiful retrospective classifier that falls apart when the assay changes.

Clinical-parameter models are useful bridges, not substitutes

There is an obvious reason to ask whether methylation subgroup can be approximated from ordinary clinical data. Not every center can obtain methylation profiling quickly, and not every health system can absorb specialized assays on demand. Imaizumi and colleagues tested machine learning models using only clinical parameters: age, fetal hemoglobin, platelet count, and PTPN11 status. In 128 training patients and 72 validation patients, decision tree, support vector machine, and naïve Bayes approaches reached 82% to 84% concordance with array-based methylation classification.[4]

That is attractive if the alternative is no methylation-informed estimate at all. It is not attractive enough to confuse with a methylation classifier. Concordance of 82% to 84% means roughly 16% to 18% of patients would be assigned differently than with array-based methylation classification.[4] In risk-adapted leukemia care, that is not a small clerical inconvenience. It can mean a child is treated as biologically lower or higher risk than the methylation reference would suggest.

ApproachPractical appealMain limitation
Genome-wide methylation arraysRich molecular profiling and discovery-grade contextLess convenient for rapid routine decision-making
124-CpG XGBoost consensus classifierValidated three-tier methylation assignment across EPIC arrays and targeted MethylSeqRetrospective validation still needs prospective multi-center confirmation
Clinical-parameter ML modelsAccessible where methylation profiling is unavailableOnly 82% to 84% concordant with array-based methylation classification
BMP4 single-locus assayFast, simplified methylation readoutA single-locus biomarker is a trade-off against full consensus classification

The clinical-parameter models are best understood as bridges. They may help centers think probabilistically when methylation testing is unavailable, delayed, or logistically impossible. They should not be presented as equivalent substitutes for a methylation classifier, because the evidence does not support that stronger claim.

The BMP4 assay sits in the middle

A different compromise is the single-locus methylation assay. Ghanjati and colleagues evaluated BMP4 methylation by bisulfite next-generation sequencing in 111 patients, reporting a 48-hour turnaround and 0.89 specificity compared with array-based methylation classification. Five-year disease-free survival was 38% in the high-methylation group and 62% in the normal-methylation group, with p=0.007.[5]

The attraction is obvious: a faster and simpler assay fits the tempo of clinical decision-making better than a workflow that may take weeks. But the trade-off is just as obvious. BMP4 is a simplified biomarker, not the full 124-CpG consensus classifier. It may be useful when speed and feasibility dominate, but it should not be inflated into the same evidentiary category as the cross-platform XGBoost system.

From stratification to treatment assignment: the TRAZA trial

The decisive test for this field is not whether methylation subgroups can be named. It is whether assigning therapy by those subgroups improves what happens to children. That is why the St. Jude TRAZA trial is the natural endpoint of the current evidence chain. Active in 2026, TRAZA is the first prospective trial to assign JMML therapy based on methylation subgroup risk stratification.[6]

In TRAZA, low-risk patients receive azacitidine plus trametinib, while high-risk patients receive the same combination plus fludarabine and cytarabine.[6] That design moves methylation classification from prognostic description into treatment allocation. It is the point at which a classifier stops being merely a way to explain past outcomes and becomes part of a prospective therapeutic strategy.

This is the right level of ambition. The consensus methylation classifier is currently the strongest prognostic framework in JMML and a serious candidate for guiding risk-adapted therapy. But prospective use is the standard that matters now. Retrospective 98% accuracy is not the same as prospective multi-center reliability, and clinical trial stratification is not the same as a cleared standalone treatment device.

For now, the defensible position is cautious: DNA methylation machine learning has made JMML risk stratification more biologically legible and more clinically useful. The field is now testing whether that framework can safely guide treatment intensity before it is called settled practice.

References

  1. The genomic landscape of juvenile myelomonocytic leukemia, Nature Communications, 2017.
  2. RAS-pathway mutation patterns define epigenetic subclasses in juvenile myelomonocytic leukemia, Nature Communications, 2017.
  3. International consensus definition of DNA methylation subgroups in juvenile myelomonocytic leukemia, Clinical Cancer Research, 2021.
  4. Machine learning approach for predicting DNA methylation classification of juvenile myelomonocytic leukemia using clinical parameters, Scientific Reports, 2022.
  5. BMP4 methylation as a simplified biomarker for juvenile myelomonocytic leukemia risk stratification, Clinical Epigenetics, 2025.
  6. TRAZA: Trametinib and Azacitidine for Juvenile Myelomonocytic Leukemia, St. Jude Children’s Research Hospital.