Skip to main content
ClinicalMind logoClinicalMind

What the evidence says about AI pancreatic cancer detection on CT

A structured comparison of three leading AI tools for pancreatic cancer detection on CT—PANORAMA, PANDA, and REDMOD—using the PROBAST+AI framework reveals that performance claims differ sharply under critical scrutiny. The strongest evidence comes from PANORAMA's international non-inferiority trial, while PANDA's real-world validation is limited to East Asian populations with disclosed conflicts, and REDMOD's pre-diagnostic advantage rests on a single-centre cohort without independent replication.

Tool
PANORAMA, PANDA, REDMOD
Manufacturer
Alibaba DAMO Academy, Mayo Clinic, EU Horizon 2020 consortium
Updated

Reviewer

Editorial Team

Editorial staff of the publication

FDA clearance status

PANDA: FDA Breakthrough Device designation; PANORAMA and REDMOD: research-stage

A regulatory fact, reported separately from the evidence verdict.

Risk-of-bias verdict

Moderate (PANORAMA) to high (PANDA, REDMOD)

The hard part in reviewing AI evidence for pancreatic cancer research is not deciding whether the reported numbers are impressive. They are. The hard part is deciding what those numbers actually apply to: detecting visible pancreatic ductal adenocarcinoma on contrast-enhanced CT, opportunistically flagging cancer on non-contrast CT, or finding a pre-diagnostic signal before the cancer is visible to a reader.

This article is an evidence appraisal, not clinical advice. It separates regulatory status from evidence strength and uses PROBAST+AI as the organizing standard for risk of bias and applicability, while acknowledging one caveat at the start: PROBAST+AI was built for prediction models using regression or artificial intelligence methods, and its use for imaging-based detection models is still less settled than for traditional tabular risk prediction models.[1]

The bottom line is plain: PANORAMA, PANDA, and REDMOD each report AI performance that exceeds radiologists on at least one detection task, but the studies are not interchangeable. PANORAMA is the strongest evidence anchor for concurrent detection on contrast-enhanced standard-of-care CT; PANDA has the most regulatory momentum and striking non-contrast CT performance, with important population and conflict-of-interest constraints; REDMOD is the most provocative pre-diagnostic signal, but also the least deployment-ready evidentiary package. No head-to-head trial compares the three tools.[2][3][4]

Radiology CT slice with pancreas region highlighted and AI-style diagnostic overlays

Three Claims That Sound Similar but Are Not

Pancreatic cancer creates unusually strong pressure to believe in earlier detection. The clinical stakes include a large share of late-stage diagnosis, poor five-year survival, and existing attention to imaging triggers in older adults with new-onset diabetes.[5] That pressure is exactly why the detection task has to be kept precise. A tool that helps a reader on a contrast-enhanced CT ordered for clinical reasons is answering a different question from a tool that screens non-contrast CTs ordered for other purposes. A tool that looks back years before diagnosis is answering a harder question again.

ToolCT settingClinical questionWhat the headline comparison can and cannot mean
PANORAMAStandard-of-care contrast-enhanced CTCan AI detect PDAC at the time of clinically interpretable imaging?Most relevant to concurrent detection support in radiology reading; not evidence for population screening.
PANDANon-contrast CTCan AI opportunistically detect pancreatic cancer on scans not optimized for pancreas evaluation?Most relevant to broad CT triage workflows; performance must be interpreted with PPV, population, and deployment setting.
REDMODPre-diagnostic CT using radiomicsCan AI detect a visually occult signal before eventual PDAC diagnosis?Most relevant to early-warning research; not yet independently replicated or prospectively proven.

A governance committee should not treat those as three versions of the same procurement claim. They differ in scan type, disease visibility, prevalence setting, reader role, and the consequence of a false positive.

The Appraisal Lens: PROBAST+AI, With Imaging Caveats

PROBAST+AI asks the right family of questions for this comparison: who was included, where the data came from, what predictors were used, how outcomes were defined, and whether the analysis could overstate performance.[1] For imaging AI, the same domains matter, but the answers often live in details that are easy to miss in a procurement slide: scanner mix, contrast phase, reader design, prevalence, case enrichment, external validation, and whether the algorithm is being tested against radiologists under fair conditions.

The practical questions are therefore narrower than “Does AI beat radiologists?” They are: Did it beat radiologists on the same scans? Were the readers representative? Was the study population close to the intended deployment population? Was the outcome ascertainment reliable? Was performance tested outside the development environment? Was the positive predictive value high enough to survive a lower-prevalence workflow? And is regulatory progress being mistaken for clinical validation?

PANORAMA: The Cleanest Test of Concurrent Detection

PANORAMA deserves the longest look because its design most closely resembles evidence a hospital committee can interrogate. The Lancet Oncology study was an international, paired, non-inferiority, confirmatory observational reader study using standard-of-care CT scans. It included 3,440 patients and 68 radiologists from 12 countries, and compared AI performance with a radiologist pool on pancreatic cancer detection.[2]

The main performance result is not just a high standalone number. PANORAMA reported AI AUROC of 0.92 versus 0.88 for the radiologist pool, with a superiority p value of 0.001.[2] That paired reader design matters because it reduces the common ambiguity in AI papers where an algorithm is tested on one task and human readers are tested under a different, sometimes less favorable, condition.

The international reader pool also matters. Sixty-eight radiologists from 12 countries do not automatically guarantee local applicability, but they give a department chair more to work with than a single-institution retrospective comparison. The study asks a concrete question: on clinically acquired CT scans, can the model detect PDAC at least as well as, and in this report better than, radiologists interpreting the same kind of material?[2]

The funding and conflict profile is also relevant. PANORAMA was funded through EU Horizon 2020, with no commercial entity identified behind it.[2] That does not make the result immune to bias, but it removes one procurement complication that appears more sharply with PANDA: the need to distinguish scientific performance from a vendor’s commercial interest.

There is still a limit to what can be independently checked from the accessible record. The full text is behind a paywall, and the pre-specified statistical analysis plan could not be verified beyond the abstract and supplementary materials available in the appraisal record.[2] That is not a reason to dismiss the study; it is a reason not to let the strongest evidence package become a frictionless deployment argument.

PANDA: Striking Performance, Sharper Procurement Questions

PANDA addresses a different and operationally tempting problem: pancreatic cancer detection on non-contrast CT. Nature Medicine reported AUCs from 0.986 to 0.996, real-world validation on 20,530 patients, and 34.1% better sensitivity than radiologists.[3] For anyone overseeing CT volumes, that is the kind of result that can quickly become a screening-workflow conversation.

The procurement tension is that the same study also gives the committee the metric it should not skip: positive predictive value. PANDA reported PPV of 56%, or 68.9% after adjustment.[3] Sensitivity gains are important, but PPV determines how many downstream workups, callbacks, specialist reviews, and patient anxieties a low-prevalence deployment could create. A tool can be technically impressive and still require careful local modeling before it belongs in a live pathway.

PANDA is also the only tool in this comparison with a documented FDA regulatory milestone. In April 2025, Targeted Oncology reported that the tool had received FDA Breakthrough Device designation.[6] That designation is meaningful because it signals agency interest in expedited development for a serious condition, but it is not FDA clearance, approval, or proof that the tool has already demonstrated clinical benefit in the intended deployment setting.

Generalisability is the next constraint. PANDA’s evidence is limited to East Asian populations, and the research record does not establish performance outside that region.[3] That limitation is not a technical footnote. Pancreatic anatomy, disease prevalence, CT acquisition patterns, referral pathways, and threshold for downstream workup may all differ across health systems. A hospital outside the validated population would be extrapolating, not simply adopting.

The conflict-of-interest profile also has to be read in the main analysis, not parked in a disclosure paragraph no one operationalizes. PANDA was developed by Alibaba DAMO Academy, and the article disclosed that authors were company employees holding stock and patents.[3] Disclosed conflicts do not invalidate a model. They do raise the standard for independent validation, local monitoring, and contract language around performance claims, post-market drift, and data use.

Three evidence-strength pillars representing different levels of medical AI validation

REDMOD: The Most Ambitious Claim Is Also the Least Settled

REDMOD is not just another detection model with a lower AUC. It is attempting something harder: radiomics-based detection of visually occult pancreatic cancer before the eventual diagnosis. Gut published the study in 2026, reporting a median 475-day lead time, AUC of 0.82, and 73% sensitivity versus 38.9% for radiologists.[4]

The “up to three years earlier” framing is clinically compelling, and it should be. If a model could reliably surface a pre-diagnostic signal while intervention is still more plausible, the value would be different from merely accelerating interpretation of an already suspicious scan. The lower AUC should not be read as simple inferiority to PANORAMA or PANDA; REDMOD is working closer to the edge of detectability, where the target may not be visually apparent to radiologists.

But that same novelty makes the evidence constraints more important. REDMOD’s training cohort was single-centre at Mayo Clinic, no independent replication exists, and the planned prospective AI-PACED trial was not registered on ClinicalTrials.gov as of the search date. The model also has no FDA regulatory milestone in the research record.[4]

The right reading is therefore narrow: REDMOD supports the possibility that CT radiomics may contain pre-diagnostic PDAC signal before routine interpretation detects cancer. It does not yet establish a deployable early-detection pathway, a stable false-positive burden, or generalizable performance across institutions.

Regulatory Status Is Not the Evidence Verdict

PANDA’s FDA Breakthrough Device designation gives it regulatory momentum that PANORAMA and REDMOD do not have.[6] That matters for development pathway, sponsor engagement, and the possibility of expedited review. It should not be converted into a claim that the device is cleared, approved, clinically effective, or ready for broad deployment.

PANORAMA and REDMOD remain research-stage in this appraisal. PANORAMA has the more governance-ready reader-study design, while REDMOD has the more exploratory early-detection claim. Neither status is a shortcut to a yes-or-no decision. A hospital still has to ask whether its patient population, CT protocols, scanner vendors, contrast use, radiologist workflow, and downstream pancreatic clinic capacity resemble the conditions in the evidence.

Scorecard for Governance Review

Appraisal questionPANORAMAPANDAREDMOD
Primary CT settingStandard-of-care contrast-enhanced CTNon-contrast CTPre-diagnostic CT radiomics
Detection taskConcurrent PDAC detectionOpportunistic detection on non-contrast CTVisually occult, pre-diagnostic signal detection
Study design strengthInternational paired non-inferiority reader studyLarge-scale deep learning study with real-world validationRecent radiomics study with longitudinal pre-diagnostic framing
Population and scale3,440 patients; 68 radiologists from 12 countriesReal-world validation on 20,530 patients; limited to East Asian populationsSingle-centre Mayo Clinic training cohort
Key performance resultAI AUROC 0.92 vs radiologist pool 0.88; superiority p=0.001AUC 0.986-0.996; 34.1% better sensitivity than radiologists; PPV 56% and 68.9% adjustedAUC 0.82; median 475-day lead time; 73% sensitivity vs 38.9% for radiologists
Regulatory statusResearch-stage in this appraisalFDA Breakthrough Device designation reported in April 2025Research-stage in this appraisal; no FDA milestone identified
Main bias or applicability concernFull text paywall limits independent verification of all protocol detailsEast Asian generalisability limits and material Alibaba DAMO Academy conflictsNo independent replication, no registered prospective trial as of search date
Most defensible interpretationStrongest evidence for concurrent CT detection supportMost regulatory momentum, but local deployment risk depends heavily on population and PPVMost provocative early-detection signal, still unconfirmed

What the Evidence Can Carry

The published evidence supports the proposition that AI can outperform radiologists in CT studies of pancreatic cancer detection. It does not support treating all three tools as interchangeable, and it does not establish that one is superior to another in practice. Without a head-to-head comparison, the safest comparison is methodological: PANORAMA has the strongest reader-study evidence for concurrent contrast-enhanced CT detection; PANDA has the most regulatory momentum and some of the highest reported performance numbers, with unresolved population and commercial-conflict concerns; REDMOD has the most clinically provocative pre-diagnostic signal, but remains unconfirmed.

That distinction is the piece worth carrying into a radiology AI review. The question is not whether the papers are exciting. It is which claim each paper actually proves, which setting it applies to, and which uncertainty would become the hospital’s responsibility after deployment.

References

  1. PROBAST+AI: an updated quality, risk of bias, and applicability assessment tool for prediction models using regression or artificial intelligence methods — BMJ, 2025.
  2. Artificial intelligence and radiologists in pancreatic cancer detection using standard of care CT scans (PANORAMA): an international, paired, non-inferiority, confirmatory, observational study — Lancet Oncology, 2025.
  3. Large-scale pancreatic cancer detection via non-contrast CT and deep learning — Nature Medicine, 2023.
  4. Next-generation AI for visually occult pancreatic cancer detection in a low-prevalence setting with longitudinal stability and multi-institutional generalisability — Gut, 2026.
  5. AI-assisted screening for pancreatic cancer — Lancet Oncology, 2025.
  6. AI Tool Earns FDA Breakthrough Device Designation in Pancreatic Cancer — Targeted Oncology.

Risk-of-bias scorecard

Study design
International paired reader study, large-scale deep learning, single-center radiomics
External / prospective validation
Included for PANORAMA (international) and PANDA (East Asian); not for REDMOD
Key performance metric
AUROC 0.92 (PANORAMA), AUC 0.986-0.996 (PANDA), AUC 0.82 (REDMOD)
Overall rating
Moderate (PANORAMA) to high (PANDA, REDMOD)

Informational only — read the full disclaimer. This content supports procurement and research judgment, not clinical care decisions.

Submit a correction or sourcing issue

Blogarama - Blog Directory