AI is beginning to answer a version of cancer recovery prognosis that clinicians can actually use: not a vague promise that software can “predict the future,” but a risk estimate that may help decide who needs closer surveillance, perioperative optimization, treatment intensification, or a different conversation before therapy starts. The strongest published results through mid-2026 are meaningful. In Stanford’s MUSK study, the model correctly predicted disease-specific survival 75% of the time across 16 cancer types, compared with 64% for standard AJCC staging-based predictions; in melanoma, it identified patients most likely to relapse within 5 years with 83% accuracy, about 12 percentage points above other foundation models.[1]
That is not a small gain. A difference of 8 to 12 percentage points can change which patients are watched more closely, which postoperative risks are discussed before surgery, and which borderline treatment decisions receive extra scrutiny. But the clinical status of these tools is more restrained than the performance numbers suggest. As of mid-2026, no AI prognosis model has FDA clearance for primary clinical decision-making, and the leading prognosis studies remain largely retrospective rather than prospective, locked-down, multi-institutional clinical validations.

The Best Evidence Is No Longer Just About Detection
Much of the public discussion around oncology AI still drifts toward detection: can the model find cancer on a slide or scan? That matters, but prognosis asks a different question. A patient already diagnosed with cancer and sitting in a treatment-planning visit needs to know whether the disease is likely to recur, whether survival risk is higher than staging alone implies, whether immunotherapy is worth the toxicity risk, or whether surgery carries a complication profile that should change preparation.
MUSK is important because it moves directly into that terrain. It is a foundation model trained to integrate pathology images and language-linked clinical information, and its disease-specific survival result is easy to compare with the method oncology teams already understand: AJCC staging. The model’s 75% survival prediction accuracy versus 64% for AJCC staging does not mean it can tell an individual patient exactly what will happen. It means that, in retrospective validation, it sorted survival outcomes better than staging alone across a broad set of cancers.[1]
The melanoma relapse result is clinically sharper. A 5-year relapse-risk estimate could affect surveillance intensity, trial discussions, and the threshold for additional review in patients whose conventional risk factors leave uncertainty. The published number—83% accuracy, about 12 percentage points higher than other foundation models—deserves attention because the endpoint is not abstract model performance; it is a risk state that can shape follow-up planning.[1]
CHIEF, the Harvard pathology foundation model published in Nature in 2024, belongs in the same conversation, but not in the same scoreboard column. It achieved about 94% accuracy in cancer detection across 11 cancer types and outperformed other state-of-the-art AI methods by up to 36%.[2] That shows how strong modern pathology foundation models have become at reading tissue signal. It does not show that CHIEF’s 94% detection accuracy is better than MUSK’s 75% survival prediction accuracy, because those numbers measure different tasks, on different datasets, against different endpoints.
That distinction is not pedantry. Detection accuracy may help a pathologist find or classify cancer. Prognosis accuracy may influence monitoring, treatment planning, and discussions about likely disease course. A model can be excellent at one and unproven at the other.
What the Main Model Families Are Actually Predicting
The evidence base is easiest to read when the models are grouped by clinical question rather than by architecture. The practical issue is not whether a model is called a foundation model, XGBoost system, or multimodal network. It is what decision its output might reasonably inform.
| Evidence Area | Representative Evidence | What It Predicts | Clinical Use It Might Support | Main Caveat |
|---|---|---|---|---|
| Survival prognosis | MUSK | Disease-specific survival across 16 cancer types; melanoma relapse within 5 years | Risk stratification, surveillance intensity, treatment-planning review | Retrospective validation; not FDA-cleared for primary decision-making |
| Multimodal survival modeling | SurvPGC | Survival from pathology imaging, genomics, and clinical data | Integrated prognosis when multiple data streams are available | Performance gains are over single-modality approaches, not a universal clinical benchmark |
| Short-term mortality | Danish pan-cancer XGBoost model | 30-day mortality in advanced cancers | Urgency, goals-of-care timing, high-risk care planning | Average precision improvement is modest and task-specific |
| Breast cancer recurrence | Systematic review of 62 studies | Recurrence risk using clinical data and, in some studies, imaging | Follow-up planning and recurrence-risk research | Heterogeneous studies; summary performance is not pooled proof |
| Immunotherapy response | Gastroesophageal cancer spatial AI analysis | Response signal from H&E slides compared with PD-L1 CPS | Treatment selection research | Reported as study-level performance, not routine clinical deployment |
| Perioperative recovery | Danish colorectal cancer surgery implementation | Risk used in surgical workflow to reduce severe postoperative complications | Preoperative optimization and complication prevention | Single-center, non-randomized before/after design |
Survival: Where Multimodal Models Start to Make Clinical Sense
Survival prediction is where the appeal of multimodal AI is strongest. Human clinicians already integrate stage, grade, performance status, laboratory values, molecular markers, treatment history, and visual pathology impressions. The hard part is not knowing that these data streams matter; it is combining them consistently across many patients and cancer types without flattening the patient into a staging label.
SurvPGC, published in npj Digital Medicine in 2025, took that integration problem directly. The model combined pathology imaging with genomics and clinical data for survival modeling and consistently outperformed single-modality approaches. Its cross-attention visualization also suggested that clinical and genomic inputs directed the model toward complementary tissue regions, a useful clue that the system was not merely adding more data but learning different kinds of prognostic signal from different sources.[3]
The caution is equally important. “Outperformed single-modality approaches” is not the same as “ready to replace a tumor board.” Multimodal models can be fragile in real clinics because the very data that make them powerful are unevenly available. A tertiary cancer center may have digitized whole-slide images, molecular profiling, structured clinical variables, and longitudinal follow-up. A community hospital may have only part of that record in machine-readable form.
Short-Term Mortality: A Different Prognosis Question
Thirty-day mortality prediction in advanced cancer is a different clinical instrument from 5-year relapse prediction. It is closer to triage: who may need urgent review, supportive-care escalation, or a more explicit discussion of treatment burden and near-term risk?
A Danish pan-cancer XGBoost study published in ESMO Real World Data and Digital Oncology in 2025 reported an average precision of 0.56 for predicting 30-day mortality in advanced cancers, compared with 0.51 for cancer-specific models. The top biomarkers were plasma albumin, white blood cell count, and lactate dehydrogenase.[4] The model’s advantage is not dramatic in isolation, but the endpoint is severe and time-sensitive. Even a modest improvement may matter if it reliably identifies patients who should not wait for the next routine appointment.
Here again, the performance metric should not be stretched beyond its task. Average precision in a short-term mortality model cannot be compared directly with MUSK’s survival accuracy or CHIEF’s cancer detection accuracy. The model is answering a narrower, urgent-care question.
Recurrence: Promising Numbers, Uneven Evidence
Recurrence prediction is where patients often feel prognosis most acutely. Treatment may be complete, scans may be clear, and yet the follow-up schedule carries a hidden question: how likely is the cancer to come back?
A 2025 systematic review in Discover Oncology examined 62 breast cancer recurrence prediction studies published from 2003 through 2023. It reported that support vector machine models achieved about 87% average performance on clinical data, while combined clinical-imaging approaches reached 95% to 95.2% accuracy. The review also noted that about 30% of early-stage breast cancer patients experience recurrence within a decade.[5]
Those figures are useful, but they should not be read as a single pooled estimate of how well AI predicts breast cancer recurrence. The included studies differed in datasets, predictors, modeling methods, endpoints, and validation strategies. A weighted or average performance summary across heterogeneous studies can show that the field has signal; it cannot by itself tell a clinic which model will work for its patients next Monday.
Treatment Response: When Prognosis Becomes a Therapy Question
Some prognosis problems are inseparable from treatment choice. In gastroesophageal cancer, an AI analysis using single-cell spatial information from H&E slides reported an AUC of 0.81 for immunotherapy response, compared with 0.65 for PD-L1 combined positive score; a combined model reached an AUC of 0.84.[6] The attraction is obvious: H&E slides are already part of routine pathology, while immunotherapy decisions are high-stakes and imperfectly predicted by existing biomarkers.
But response prediction is not the same as survival prediction, and a biomarker-comparison AUC is not the same as a clinical recommendation. Before a model like this can guide treatment selection, clinicians need to know whether it improves decisions prospectively, whether it generalizes across institutions and staining workflows, and how its output should be weighed against PD-L1, molecular findings, performance status, comorbidities, and patient preference.
The One Place AI Has Moved Into a Full Clinical Workflow
The Danish colorectal cancer surgery study deserves a different kind of attention because it is not just another retrospective benchmark. Published in Nature Medicine in 2025, it described a full-scale clinical implementation in which an AI-enabled workflow was associated with a reduction in severe postoperative complications from 28.0% to 19.1%, with an odds ratio of 0.63, and about $2,848 in cost savings per patient. The study included 18,403 patients.[7]
This is not cancer prognosis in the narrow sense of long-term survival prediction. It is recovery prognosis in a very practical sense: which surgical patients are at high risk of a severe postoperative course, and what happens when that risk estimate is placed into a care pathway rather than left in a paper?
The outcome matters because postoperative complications are not secondary details. They affect mortality risk, adjuvant therapy timing, length of stay, patient function, and cost. A model that helps reduce severe complications can change the lived trajectory of cancer recovery even if it does not directly predict tumor biology.
The design also limits the claim. The Danish study was single-center and non-randomized, using a before/after comparison. That means the observed improvement cannot be treated as definitive causal proof that the AI system alone produced the reduction. Other workflow changes, secular trends, staffing patterns, or perioperative practice shifts could have contributed. A multicenter randomized trial, NCT06645015, is ongoing and should provide stronger evidence about causality and generalizability.[7]

Why Retrospective Accuracy Is Not the Same as Clinical Readiness
Retrospective validation is necessary. It is also where many oncology AI tools stop. A model can perform well on stored slides, curated clinical variables, and known outcomes, then weaken when it meets missing data, scanner variation, changing treatment protocols, demographic shifts, and local documentation habits.
For cancer recovery prognosis, the deployment threshold should be higher than leaderboard performance because the output can change behavior. A high-risk label may prompt extra surveillance, more aggressive supportive care, treatment delay, referral to another specialist, or a difficult goals-of-care conversation. A low-risk label can also cause harm if it falsely reassures a team and reduces monitoring.
The missing pieces are familiar but not optional: locked-down models, prospective validation, multi-institutional testing, predefined endpoints, subgroup performance analysis, workflow integration, and regulatory review. Ruijiang Li, the Stanford researcher behind MUSK, estimated in March 2026 that AI prognosis tools may still need 3 to 5 years before they are ready for prospective validation and clinical deployment, citing the need for locked-down models, multi-institutional validation, and FDA clearance.[8] That timeline is an expert judgment, not a validated forecast, but the requirements are the right ones.
Explainability is another practical constraint. In one survey cited in the same expert context, 85% of oncologists agreed they should be able to explain AI models, and 91% said AI developers bear responsibility for harm.[8] Those attitudes do not settle the legal question, but they show the clinical discomfort with black-box risk scores that no one can adequately justify when a patient asks why a plan changed.
For broader evidence-quality context across medical AI, this is the same gap seen in many diagnostic tools: retrospective success arrives before prospective proof. The difference in oncology prognosis is that the consequences often unfold over months or years, making harm harder to attribute and benefit harder to verify. Readers comparing this field with other AI diagnostics may find the evidence-quality framing in How Strong Is the Evidence for AI in Medical Diagnostics? useful.
What Would Make a Prognosis Model Clinically Trustworthy?
The standard should depend on the decision. A model used to flag a chart for extra review does not need the same evidentiary threshold as a model used to recommend against treatment. Still, several requirements apply across use cases.
- The model should be locked before prospective testing, so performance is not quietly improved after outcomes are known.
- Validation should include multiple institutions, scanners, pathology workflows, EHR systems, and patient populations.
- The endpoint should match the clinical action: survival, recurrence, 30-day mortality, treatment response, and postoperative complications are not interchangeable.
- Performance should be reported by relevant subgroups, not only as an overall average.
- The clinical workflow should specify who sees the prediction, when they see it, what action is expected, and who is responsible if the recommendation is wrong.
- Regulatory status should be explicit, especially if the model is being used for primary clinical decision support.
The FDA question is not a paperwork afterthought. If a prognosis model influences treatment, surveillance, or perioperative decisions, it enters the same accountability space as other clinical decision-support tools. For readers tracking that boundary, the FDA’s 2026 CDS guidance for AI clinical decision support is the relevant regulatory companion.
So, Can AI Accurately Predict Cancer Prognosis and Recovery?
The most defensible answer is yes, within defined research tasks, and not yet as a routine primary clinical decision tool. MUSK’s survival and melanoma relapse results show that AI can extract prognostic signal beyond standard staging. SurvPGC shows the logic of combining pathology, genomics, and clinical variables. The Danish short-term mortality model shows that advanced-cancer prognosis can be framed around urgent near-term risk. Breast cancer recurrence studies show a broad but heterogeneous literature. The gastroesophageal immunotherapy work shows that AI may find treatment-response signal in ordinary pathology slides. The Danish colorectal surgery deployment shows that AI-supported risk workflows can move beyond retrospective testing and be associated with better recovery outcomes.
None of those findings should be collapsed into a single claim that AI has solved cancer prognosis. They point to a narrower and more useful conclusion: AI cancer prognosis is no longer speculative as a research capability, and the best models are beginning to outperform standard methods by clinically meaningful margins. As of Q3 2026, however, most prognosis use cases remain pre-deployment until prospective, multi-institutional validation, locked model testing, workflow accountability, and regulatory clearance catch up.
References
- A vision-language foundation model for precision oncology, Nature, Jan. 2025.
- New AI tool can diagnose cancer, guide treatment, predict patient outcomes, Harvard Gazette, Sept. 2024.
- SurvPGC: multimodal deep learning for survival prediction integrating pathology, genomics and clinical data, npj Digital Medicine, Jan. 2025.
- Pan-cancer XGBoost model for 30-day mortality prediction in advanced cancers, ESMO Daily Reporter / ESMO Real World Data and Digital Oncology, July 2025.
- Machine learning models for breast cancer recurrence prediction: a systematic review, Discover Oncology, Feb. 2025.
- Single-cell spatial AI analysis of H&E slides for immunotherapy response in gastroesophageal cancer, ASCO, 2024.
- Artificial intelligence-guided perioperative care for colorectal cancer surgery, Nature Medicine, Sept. 2025.
- AI Prognosis Tools Still Need Prospective Validation and FDA Clearance, Expert Says, The ASCO Post, March 2026.
Comments
Join the discussion with an anonymous comment.