“AI improves cancer detection” is where the evidence discussion should slow down, not speed up. Which cancer? Which task? Detection by whom, at what point in the workflow? And which endpoint: more lesions found, more clinically important disease found, fewer interval cancers, lower mortality, less workload, fewer recalls, or some net patient benefit? The current evidence for AI in cancer care is strongest in two places: breast screening detection and colonoscopy adenoma detection. Even there, much of the proof sits on surrogate or intermediate endpoints rather than final patient outcomes. A systematic review and meta-analysis in the Journal of the American College of Radiology found 49 randomized controlled trials across seven cancer types, with the largest evidence base in colorectal screening and much thinner RCT evidence elsewhere. It also found that none of the included RCTs reported patient-centered outcomes. [1]
FDA clearance belongs in a different lane. It may tell a hospital that a device has passed a regulatory pathway for marketing, but it does not by itself show that a cancer-AI tool improves clinical outcomes in deployed screening. A JAMA Network Open review of 950 FDA-authorized AI/ML-enabled devices found that 97% went through the 510(k) pathway; among 717 radiology devices, only 33, or 5%, had been prospectively tested. [2]

| Evidence question | Best-supported current answer | What it does not prove |
|---|---|---|
| Where is randomized or prospective evidence most concentrated? | Breast screening detection and colonoscopy adenoma detection. [1] | That all cancer-AI use cases have comparable clinical proof. |
| What is strongest in colorectal screening? | AI-assisted colonoscopy increases adenoma detection and polyp detection in pooled RCTs. [1] | That it increases advanced adenoma detection, colorectal cancer detection, or patient-centered outcomes. |
| What is strongest in breast screening? | MASAI provides randomized evidence that AI-supported mammography can preserve interval-cancer safety while increasing sensitivity; PRAIM adds a large real-world implementation signal. [3][4][5] | That higher sensitivity automatically means lower mortality or no overdiagnosis trade-off. |
| What about lung, prostate, gastric, liver, and other cancer tasks? | The JACR RCT map shows mostly single-study signals, null findings, or absent patient-centered endpoints. [1] | A general claim that AI improves cancer care across oncology. |
| How should FDA clearance be read? | As authorization evidence, not as a substitute for prospective clinical testing or endpoint-specific benefit. [2] | That a cleared tool has proved better outcomes in the intended clinical setting. |
The RCT map is encouraging only if the endpoint is kept in view
The JACR review is the useful starting point because it does not let “AI in cancer care” behave like one evidence category. It pooled randomized trials across cancer types and separated what was actually measured. That matters because a positive detection endpoint can be clinically meaningful without being the same thing as cancer prevention, mortality reduction, or net benefit.
Colorectal screening dominates the RCT base. In 36 RCTs, AI-assisted colonoscopy increased adenoma detection with a pooled relative risk of 1.22, with a 95% confidence interval from 1.17 to 1.28. In 28 RCTs, it increased polyp detection with a pooled relative risk of 1.20, with a 95% confidence interval from 1.14 to 1.26. [1]
Those are not trivial findings. Adenoma detection rate is a familiar colonoscopy quality metric, and an AI tool that reliably increases adenoma detection in real procedures has a more concrete evidentiary footing than a model that performs well only on held-out images. But the same meta-analysis found no significant effect on advanced adenoma detection, with a relative risk of 1.07, and no significant effect on colorectal cancer detection, with a relative risk of 0.97. [1]
That split is the whole appraisal problem. A hospital can reasonably see AI-assisted colonoscopy as having randomized evidence for increasing adenoma and polyp detection. It should not silently translate that into proof that the same product prevents colorectal cancer, reduces interval cancer, or improves survival. Those may be plausible downstream aims, but they were not demonstrated in the RCT endpoint map summarized here.
The review’s risk-of-bias finding also needs a careful reading. Ninety-eight percent of the included RCTs were rated as having “some concerns” using ROB2. [1] That should make governance teams cautious, but it is not the same as saying the entire RCT base is unusable. Short-follow-up diagnostic RCTs often strain general risk-of-bias tools: blinding can be difficult, behavior changes are part of the intervention, and downstream outcomes may not be observable inside the trial window. The right conclusion is narrower: the RCT evidence is real, strongest for certain detection endpoints, and still not designed to carry patient-outcome claims.

Breast screening has the strongest prospective signal, with different caveats
Breast screening is the other area where the evidence deserves more than a dismissive “AI hype” label. The MASAI randomized trial enrolled 105,934 women and tested AI-supported mammography screening against standard screening. Its interval cancer rate was non-inferior: 1.55 per 1,000 in the AI-supported group versus 1.76 per 1,000 in the control group, with a ratio of 0.88 and a 95% confidence interval from 0.65 to 1.18. Sensitivity was higher with AI-supported screening, 80.5% versus 73.8%, while specificity was identical at 98.5%. [3][4]
Those are unusually useful results for a screening AI appraisal because they connect the detection claim to interval cancers rather than stopping at reader-level accuracy. MASAI also reported fewer unfavorable interval cancers in the AI-supported arm, including fewer invasive interval cancers, fewer T2 or larger interval cancers, and fewer non-luminal A interval cancers. [4] Still, interval cancer is not mortality. Higher sensitivity is not automatically a net patient benefit if it also changes overdiagnosis, follow-up burden, or treatment pathways in ways the trial was not powered or long enough to settle.
PRAIM is a different kind of evidence. It was a large real-world implementation study in Germany, not a randomized trial. Across 463,094 women at 12 sites, AI-supported double reading was associated with a 17.6% higher breast cancer detection rate, 6.7 versus 5.7 cancers per 1,000 screened women, with a 95% confidence interval from +5.7% to +30.8%. Recall did not worsen; the reported change was −2.5%. [5]
That is a serious implementation signal because it comes from routine screening rather than a laboratory dataset. But its design sets the ceiling on the claim. PRAIM was observational, vendor-funded by Vara, and used propensity-score adjustment to address reading-behavior bias. [5] Those features do not erase the result, but they keep it below MASAI’s randomized footing. The unresolved overdiagnosis question matters too: increased detection, including more ductal carcinoma in situ, can be helpful, harmful, or mixed depending on which additional cancers would have become clinically important.
The practical breast-screening conclusion is not “deploy every mammography AI tool.” It is that breast screening has moved beyond purely retrospective reader studies. For departments already tracking interval cancers, recalls, sensitivity, specificity, arbitration workload, and cancer characteristics, MASAI and PRAIM give a defensible evidence base to discuss a specific screening workflow. A broader breast-only discussion belongs in a separate AI cancer screening evidence review; the point here is endpoint discipline.
Outside the two clusters, the RCT evidence thins quickly
Once breast screening and colorectal endoscopy are separated out, the randomized cancer-AI evidence map becomes much sparser. The useful way to read the remaining findings is not cancer burden by cancer burden, but endpoint by endpoint: what was tested, whether it moved, and whether the trial tells the people deploying the tool anything about patient-centered benefit.
| Cancer or task area | RCT signal reported in the JACR review | Appraisal reading |
|---|---|---|
| Breast cancer detection | Relative risk 1.20, 95% CI 1.00–1.45. [1] | A positive or borderline detection signal, stronger when read alongside MASAI, but still endpoint-specific. |
| Actionable lung nodules | Relative risk 2.38, 95% CI 1.25–4.55. [1] | A single-task signal for finding actionable nodules, not proof of better lung cancer outcomes. |
| Prostate cancer | Relative risk 1.40, 95% CI 1.10–1.77. [1] | A promising detection finding, but not enough to generalize across prostate screening, diagnosis, or treatment decisions. |
| Gastric cancer | Effect not statistically significant. [1] | No RCT-supported improvement claim from this evidence map. |
| Liver cancer | Effect not statistically significant. [1] | No RCT-supported improvement claim from this evidence map. |
| Patient-centered outcomes across included RCTs | Zero included RCTs reported patient-centered outcomes. [1] | The map cannot establish survival, quality-of-life, or net-benefit claims. |
The prostate signal is a good example of why a cancer-by-cancer map is more useful than an “AI works” verdict. A positive detection result may justify further evaluation of a defined tool in a defined pathway, but prostate cancer appraisal has to keep separate PSA triage, MRI interpretation, biopsy targeting, grade prediction, treatment selection, and active surveillance. Those are different decisions with different harm profiles. A narrower discussion of that screening problem is handled in the AI prostate cancer screening evidence appraisal.
The same restraint applies to lung nodules. A higher rate of actionable nodule detection is not a trivial endpoint; missed actionable nodules can matter. But the phrase “actionable lung nodules” already signals that the trial measured a workflow-adjacent detection endpoint, not lung cancer mortality, avoided invasive procedures, anxiety, downstream imaging burden, or net benefit.
Why model-development literature should not be read like MASAI or colonoscopy RCTs
A large share of oncology AI literature lives outside randomized and prospective implementation evidence. That does not make it useless. Model-development studies can identify promising signals, improve technical methods, and support future trials. They just should not be promoted as if they had already survived clinical deployment.
Risk-of-bias tools for AI prediction studies exist precisely because apparent performance can travel poorly across institutions, scanners, staining protocols, patient mix, and workflow. PROBAST+AI was developed to extend prediction-model risk-of-bias and applicability assessment to artificial intelligence studies. [6] The need for that kind of tool is not academic nitpicking; it is what separates an internally impressive model from one that can be trusted enough to change care.
Pathology illustrates the gap. A scoping review of lung pathology AI found that about 10% of 239 development papers had external validation, and 86% of validation studies were at high risk of bias in participant selection. [7] That is a very different evidentiary posture from a randomized screening trial. It may support research enthusiasm, but it does not support broad procurement claims about clinical benefit.
An oncology-wide validation scoping review points in the same direction. It identified 56 externally validated clinically useful AI models from 2014 through 2022, and 49 of those validations were retrospective. [8] External validation is better than no external validation, but retrospective validation is still a long way from knowing what happens when the model is inserted into a clinic, a tumor board, a reading room, or a pathology queue.
This is why broad treatment-optimization, survivorship, and pathology claims should be appraised on their own evidence rather than pulled upward by breast screening and colonoscopy results. A pathology model, a radiology triage tool, a survivorship risk predictor, and an endoscopy overlay do not share one endpoint just because all are labeled AI. The evidence may be promising in one lane and absent in another.

FDA clearance has to be read against the evidence map
Clearance is often where endpoint slippage becomes operational. A vendor slide says the device is FDA cleared; the next slide says it improves cancer detection; the room quietly hears clinical benefit. Those are three different statements. The regulatory fact may be true, the detection claim may be true for a particular endpoint, and the clinical-benefit claim may still be unproved.
The JAMA Network Open clearance review is useful because it gives scale to the problem. In a market of 950 AI/ML-enabled medical devices, nearly all were cleared through 510(k), and only a small minority of radiology devices had prospective testing. [2] That does not mean cleared devices are unsafe or ineffective. It means clearance volume should not be mistaken for a body of prospective clinical trials.
For procurement and governance teams, the evidence question therefore starts after the clearance question. The next issue is whether the authorized clinical task, intended workflow, endpoint, and study design match the claim being made. Device-count context can help frame the market, but it cannot answer those deployment questions by itself; a separate AI in healthcare statistics tracker is not a substitute for endpoint review.
The evidence verdict
AI in cancer care has credible, prospective evidence in narrow detection tasks. Breast screening has the strongest prospective screening signal, especially MASAI, with PRAIM adding a large real-world but observational implementation result. Colonoscopy AI has a substantial RCT base for adenoma and polyp detection.
The same evidence does not support a broad claim that AI improves cancer outcomes. In colorectal screening, the RCT signal is strong for adenoma detection and polyp detection, but not for advanced adenoma detection or colorectal cancer detection in the JACR meta-analysis. In breast screening, sensitivity and interval-cancer findings are important, but mortality and overdiagnosis remain separate questions. In lung, prostate, gastric, liver, pathology, treatment optimization, and survivorship, the evidence is either much thinner, endpoint-specific, retrospective, high-risk-of-bias, or outside the patient-centered outcome frame.
So the defensible answer is conditional: AI can improve cancer detection for particular tasks and endpoints, with the strongest evidence in breast screening and colonoscopy adenoma detection. It has not, in the RCT map cited here, proved broad patient-centered benefit across cancer care. FDA clearance should be read beside that map, not in place of it.
References
- Artificial Intelligence for Cancer Detection: A Systematic Review and Meta-Analysis of Randomized Controlled Trials — Journal of the American College of Radiology, 2025.
- Clinical Validation and Regulatory Pathways of AI/ML-Enabled Medical Devices — JAMA Network Open.
- Screening mammography with artificial intelligence and radiologists versus radiologists alone (MASAI): a randomised, controlled, non-inferiority trial — The Lancet, 2025.
- Randomized Trial Shows AI-Supported Mammography Improves Sensitivity and Lowers Interval Cancer Rate — The ASCO Post, February 2026.
- Nationwide real-world implementation of artificial intelligence for cancer detection in population-based mammography screening — Nature Medicine, 2025.
- PROBAST+AI: a risk of bias and applicability assessment tool for prediction model studies in artificial intelligence — BMJ, 2025.
- External validation of artificial intelligence models in lung pathology: a scoping review — PMC.
- External validation of artificial intelligence models for clinical oncology: a scoping review — BMC Medical Research Methodology, 2025.