AI in cancer research has moved well past the demonstration phase in several screening and diagnostic settings. The harder question is no longer whether an algorithm can find a suspicious lesion, cell cluster, or methylation signal. It is whether that finding improves a clinical decision in a defined population, against a fair comparator, without pushing patients into avoidable biopsies, resections, callbacks, or anxiety.
The evidence now looks strongest where AI is attached to an existing screening workflow with measurable handoffs: mammography reads, lung nodule classification on low-dose CT, and selected pathology review tasks. It is more uneven where the promise is broader, earlier, or less tethered to a familiar decision point, as with multi-cancer liquid biopsy and general-purpose pathology models.

| Modality | What the evidence shows | Closest clinical question | Unresolved adoption concern |
|---|---|---|---|
| Mammography | A 2025 study of 463,094 women screened by 119 radiologists reported a 17.6% higher cancer detection rate with AI-supported screening. | Can AI increase cancer detection in population breast screening without overwhelming recall and review workflows? | Detection gain still has to be interpreted alongside recalls, interval cancers, workload, and local reading practice. |
| Lung LDCT | AI lung nodule assessment has reported AUC values around 0.93-0.95, and C-Lung-RADS reported 87.1% sensitivity versus 63.3% for standard Lung-RADS. | Can AI improve nodule risk classification inside an established LDCT screening pathway? | Higher sensitivity needs integration with Lung-RADS categories, follow-up thresholds, and false-positive management. |
| Digital pathology | LYNA reached 99% accuracy and improved pathologist micrometastasis detection from 83.3% to 91.2%; CHIEF reported AUROC 0.9397 across 15 datasets and 11 cancer types. | Can AI reduce missed findings and support triage or review in high-volume pathology? | Generalizability depends on scanner, staining, tissue handling, population mix, and external validation. |
| Liquid biopsy | GRAIL Galleri reported 99.5% specificity across more than 50 cancer types, but sensitivity was stage-dependent, from 16.8% at stage I to 90.1% at stage IV in a 4,077-participant study. | Can a blood test detect clinically meaningful cancers early enough to change outcomes? | High specificity does not solve low early-stage sensitivity, diagnostic workup burden, or outcome uncertainty. |
| Colorectal CADe | A 2025 meta-analysis of 21 RCTs and 18,232 patients found increased adenoma detection but no improvement in advanced adenoma detection, with more non-neoplastic polyp removal. | Does more detection during colonoscopy identify lesions that matter? | More removals can mean more intervention without a proportional gain in clinically important findings. |
Mammography Is the Clearest Case for Workflow-Level Evidence
Breast screening is where the current evidence feels most operational rather than theatrical. In the Eisemann et al. 2025 study, AI-supported mammography was evaluated across 463,094 women screened by 119 radiologists and was associated with a 17.6% higher cancer detection rate.[1] That is the kind of evidence administrators and screening leads can actually argue over, because it sits near the daily mechanics of a program: image interpretation, callback pressure, radiologist capacity, and cancer yield.

The important feature is not simply that the model detected more cancers. It is that the result came from a large screening context with many radiologists, rather than a narrow retrospective contest between an algorithm and a curated image set. Screening mammography already accepts a controlled amount of false-positive investigation in exchange for earlier detection. An AI tool that shifts that balance has to be judged by what happens to the whole pathway: second reads, arbitration, recall rates, biopsy load, interval cancers, and the distribution of detected disease.
A 17.6% higher detection rate is not a standalone adoption decision. It is a reason to look closely at deployment conditions. If the gain comes with manageable recall, reliable performance across breast density and age groups, and no hidden penalty for underrepresented patients, it becomes clinically persuasive. If it depends on local prevalence, reader behavior, or a population unlike the one being screened, the number travels less well.
This distinction matters because mammography AI is often sold as relief for radiologist workload. In practice, a useful system may initially increase work in specific parts of the pathway. More suspicious findings mean more adjudication, more patient communication, and more downstream procedures. A screening program should therefore evaluate the model not only at the image level, but at the level where patients experience the system: whether the callback was necessary, whether the cancer would otherwise have been missed, and whether earlier detection changed management.
Lung LDCT Has Strong Signals, but the Comparator Is the Pathway
Low-dose CT lung cancer screening gives AI a different target. The central problem is not only finding nodules; it is deciding which nodules deserve surveillance, diagnostic imaging, biopsy, or surgical evaluation. Venkadesh et al. reported AUC values in the 0.93-0.95 range for AI-based lung cancer screening assessment, and C-Lung-RADS reported 87.1% sensitivity compared with 63.3% for standard Lung-RADS.[2]
That sensitivity gain is hard to ignore. Lung screening programs are already built around risk stratification and follow-up intervals, so an AI system that better identifies malignant nodules could reduce missed cancers or accelerate workup for patients who need it. The model output, however, does not arrive in a vacuum. It lands inside Lung-RADS categories, radiology reports, multidisciplinary nodule clinics, and a patient population with competing risks from smoking-related disease.
The adoption question is therefore narrower than “Can AI detect lung cancer?” A useful LDCT tool must show whether it changes nodule classification in a way that improves care. If sensitivity rises by moving many patients into more intensive follow-up, clinicians need to know the false-positive cost. If AI identifies subtle malignant patterns that standard categorization misses, the program needs a protocol for who reviews the finding, when follow-up changes, and how disagreement between the algorithm and the radiologist is resolved.
This is where lung screening resembles mammography: the most credible AI evidence is tied to an existing bottleneck, but deployment still depends on the surrounding clinical machinery. A strong AUC does not tell a nurse navigator which patient to call first, or tell a thoracic surgeon whether a biopsy threshold should move. Those decisions need prospective evidence and local workflow testing.
Pathology AI Is Clinically Serious, but Unevenly Mature
Digital pathology is one of the most plausible homes for diagnostic AI because the clinical task is often visual, repetitive, and high-stakes. The strongest examples are not vague claims about replacing pathologists. They are assistance tasks: flagging suspicious regions, prioritizing slides, reducing oversight errors, or standardizing review in settings where subspecialty expertise is limited.
LYNA, a lymph node metastasis detection system, reached 99% accuracy and improved pathologist detection of micrometastases from 83.3% to 91.2%, with p=0.023.[3] That is a clinically legible gain. A missed micrometastasis can affect staging and adjuvant treatment discussions, while an AI-highlighted region can be verified by a pathologist rather than accepted blindly.
Paige Prostate is also important because it represents an FDA-cleared digital pathology application for prostate cancer.[3] Clearance does not mean every hospital should deploy the system tomorrow, but it does move the discussion from benchmark enthusiasm to procurement, integration, scanner compatibility, quality assurance, and responsibility for final sign-out.
The broader pathology direction is represented by CHIEF, a deep learning system that reported AUROC 0.9397 across 15 datasets and 11 cancer types in npj Precision Oncology in 2026.[4] A result like that points toward multi-cancer pathology intelligence rather than a single narrow detector. It is also exactly where caution should sharpen. Performance across many datasets is encouraging, but hospitals vary in scanners, staining protocols, fixation, case mix, and artifact patterns. The farther a model moves from a bounded task to a general pathology assistant, the more external validation matters.
Brain tumor AI raises related but distinct questions because intraoperative diagnosis, molecular classification, and neuro-oncology treatment planning have their own timing pressures. Readers focused on that setting may want the separate discussions of AI applications in brain oncology and real-time brain cancer diagnosis during surgery; duplicating that field here would blur rather than clarify the screening and diagnostic comparison.
Liquid Biopsy Shows Why Specificity Is Not the Whole Story
Multi-cancer early detection is attractive because it promises a single blood draw that can detect cancers without established screening programs. GRAIL Galleri reported 99.5% specificity across more than 50 cancer types in a study of 4,077 participants.[5] For a population-level test, that specificity matters. Even a small false-positive rate can send many healthy people into imaging, endoscopy, specialist visits, or invasive diagnostic procedures.
The sensitivity profile is the constraint. In the reported data, sensitivity was 16.8% for stage I disease and 90.1% for stage IV disease.[5] That does not make the test unimportant, but it changes the claim. A test that is much more sensitive in advanced disease than early disease cannot be discussed as if it uniformly solves early detection.
For clinicians, the key issue is what happens after a positive signal. A cancer signal of origin may guide the diagnostic workup, but patients still need confirmatory testing. Health systems need to know how many scans, procedures, and specialist referrals a positive blood test triggers, and whether those steps produce earlier, treatable diagnoses rather than prolonged diagnostic uncertainty.
Liquid biopsy is therefore one of the most promising and easiest-to-overstate areas of AI-enabled cancer diagnostics. High specificity supports serious evaluation. Stage-dependent sensitivity requires restraint in how the result is presented to patients.
Colorectal CADe Is the Cautionary Example: More Findings Are Not Always Better Screening

Computer-aided detection during colonoscopy gives the cleanest warning against treating detection as the endpoint. In a 2025 meta-analysis of 21 randomized controlled trials including 18,232 patients, CADe increased adenoma detection but did not improve advanced adenoma detection, and it was associated with higher rates of non-neoplastic polyp removal.[6]
That result is not a failure of AI in any simple sense. Adenoma detection rate is a familiar colonoscopy quality measure, and finding more adenomas can matter. But advanced adenomas are closer to the lesions screening programs most urgently want to identify. If AI increases removal of lesions that are non-neoplastic or unlikely to change management, the procedure becomes busier without necessarily becoming more beneficial.
The burden is not abstract. A highlighted polyp must be inspected, removed or deliberately left alone, documented, and explained. Extra resections add pathology work and may alter surveillance intervals. The patient bears the procedural risk and the consequences of being labeled with a finding. A model that increases the number of things seen has to prove that enough of those things matter.
The Evidence Gap Is Often About People, Not Algorithms

External validation is not a paperwork exercise. It is the difference between a model that works on familiar data and a model that can be trusted with patients whose imaging equipment, ancestry, comorbidities, tumor biology, and care access differ from the development cohort.
The ancestry imbalance in major cancer and genomics resources makes that concern concrete. The Cancer Genome Atlas repository has been reported as having a median 83% European ancestry representation, and the GWAS catalog has been reported as containing approximately 95% European data.[7] Models built from or benchmarked against such resources may still perform well, but the burden of proof should rise when they are deployed in populations that were underrepresented in training or validation.
This is especially important for tools that combine imaging, pathology, genomic, or blood-based signals. A mammography model may be sensitive to breast density distribution, scanner vendor, and local screening intervals. A pathology model may be sensitive to tissue preparation and staining. A liquid biopsy classifier may reflect biological and technical patterns that were not evenly represented across ancestry groups. Fairness testing has to be modality-specific rather than treated as a generic appendix.
For readers evaluating models after a diagnosis has already been made, the validation questions become somewhat different: calibration, prognosis, treatment selection, and outcome prediction. Those issues are covered more directly in how to evaluate AI prognostic models for cancer care and AI cancer prognosis and recovery. Screening AI should be held to its own standard: it acts earlier, often on asymptomatic people, and can expose many more patients to downstream workup.
What Counts as Sufficient Evidence for Deployment?
The most adoption-ready cancer AI tools are not necessarily the ones with the most elegant model architecture. They are the ones that answer a clinical question already under strain: too many mammograms for too few readers, too many lung nodules with uncertain risk, too many pathology slides where a small missed focus can affect staging.
A health system considering deployment should separate several questions that are often collapsed into one performance claim:
- Was the model tested prospectively or only retrospectively?
- Was the comparator average clinical practice, generalist readers, or subspecialist experts?
- Does the model improve a decision that changes patient management?
- What happens to recalls, biopsies, resections, follow-up imaging, and pathology volume?
- Was performance tested across the population that will actually receive the test?
- Who is responsible when the model and clinician disagree?
Prospective randomized trials remain limited for many cancer AI tools, especially outside narrow screening or diagnostic tasks. Retrospective studies can be valuable, particularly when datasets are large and external, but they cannot fully show how clinicians change behavior when an AI mark appears on the screen. They also cannot reliably measure whether patients experience better outcomes or simply more testing.
Regulatory status helps but does not settle local adoption. A cleared tool still has to fit the institution’s scanners, reporting system, staffing model, tumor board expectations, quality metrics, and patient population. The broader treatment pipeline is discussed separately in how AI is transforming cancer treatment in 2026; screening and diagnosis create their own adoption problem because they sit at the front door of cancer care.
A Graded Clinical View
Mammography and lung LDCT are closest to evidence-supported deployment because the available data connect AI performance to recognizable screening workflows. The Eisemann mammography result is compelling because of its scale and real-world radiologist involvement. The lung LDCT evidence is compelling because it addresses nodule risk assessment inside an established screening structure.
Digital pathology is clinically important and likely to expand, especially for review assistance, triage, and missed-focus detection. Its maturity varies by task. A lymph node micrometastasis detector, an FDA-cleared prostate tool, and a broad multi-cancer pathology model should not be treated as the same kind of evidence.
Liquid biopsy deserves serious attention but careful language. High specificity across many cancers is valuable, yet low stage I sensitivity means it should not be presented as a uniform early-cancer solution. Colorectal CADe shows the opposite side of the same problem: a tool can increase detection while leaving the more clinically consequential endpoint unchanged.
The next useful evidence will not be another isolated headline accuracy number. It will be prospective validation in the populations being served, fairness testing that is visible rather than assumed, and proof that AI changes patient-relevant decisions rather than merely increasing the volume of findings.
References
- Eye on AI: Applying Artificial Intelligence to Drive Cancer Research, Part 2. AACR Blog. August 18, 2025.
- Artificial intelligence in lung cancer screening. Frontiers in Oncology. 2025.
- Artificial intelligence in cancer diagnosis and pathology. Molecular Cancer.
- CHIEF deep learning system for multi-cancer pathology analysis. npj Precision Oncology. 2026.
- Multi-cancer early detection using liquid biopsy. Molecular Cancer. 2025.
- Artificial intelligence and computer-aided detection in colorectal cancer screening. Molecular Cancer.
- AI Cancer. Cancer Research Institute.
Comments
Join the discussion with an anonymous comment.