Skip to main content
ClinicalMind logoClinicalMind

AI Colorectal Cancer Diagnosis Evidence: Three Key Tensions

An institution evaluating AI-assisted colonoscopy for colorectal cancer screening faces three evidence tensions that prevent a straightforward adoption decision: the gap between RCT efficacy and real-world effectiveness, the clinical implications of increased diminutive polyp detection, and the emerging risk of diagnostic deskilling. This appraisal unpacks each tension using the latest peer-reviewed studies and shows why no single evidence stream alone justifies a confident decision.

Tool
CADe Colonoscopy Systems
Updated

Reviewer

Editorial Team

Evidence Appraisals Editorial Staff

FDA clearance status

FDA cleared (De Novo and 510(k))

A regulatory fact, reported separately from the evidence verdict.

Risk-of-bias verdict

Moderate

The practical question in AI colorectal cancer diagnosis evidence is not whether computer-aided detection can make a polyp box appear on a colonoscopy screen. FDA-cleared CADe products exist, and randomized trials have repeatedly shown that endoscopists detect more adenomas when the software is running. The harder question is whether a hospital, ambulatory endoscopy network, or health system should treat that trial signal as enough to justify broad deployment.

At the committee table, the evidence does not arrive as one clean answer. It arrives as three streams that do not fully converge: RCT efficacy is stronger than real-world effectiveness; the additional lesions are often diminutive adenomas whose downstream consequences matter; and the same second-look function that may reduce miss rates may also create a governance problem if unaided skill weakens over time.

Appraisal fieldCurrent evidence status
Tool categoryReal-time computer-aided detection during colonoscopy, usually flagging suspected polyps for the endoscopist’s review
Relevant cleared productsGI Genius, SKOUT, DISCOVERY, ENDO-AID, and CAD EYE are listed in the FDA-cleared CADe landscape described in the available source record; clearance should be read as market authorization, not proof of local clinical value.[1]
Study designs representedRandomized controlled trials and meta-analyses, pragmatic implementation studies, tandem miss-rate studies, economic models, and one early observational deskilling signal
Most consistent positive signalHigher adenoma detection in RCT settings, with relative-risk estimates in meta-analyses generally in the 1.24–1.78 range and an all-RCT 2024 meta-analysis covering more than 18,000 participants.[2]
Main institutional concernBenefits seen under trial conditions have not reliably reproduced in normal workflow, and much of the gain is concentrated in adenomas smaller than 5 mm.
Adoption postureReasonable to evaluate conditionally; not yet settled enough for universal deployment as a standard purchase.
Three diverging AI colonoscopy evidence pathways at a decision crossroads

The RCT Signal Is Real, but It Is Not the Same as Implementation

The randomized evidence deserves respect. Across RCT meta-analyses summarized in the available review literature, CADe increases adenoma detection rates, with reported relative-risk improvements of roughly 20% to 50% depending on the included trials and analytic approach.[2] That is not a trivial finding. If an endoscopist misses a lesion that the system would have highlighted, the distinction between “statistically significant” and “operationally inconvenient” will not comfort the patient.

But institutional adoption is usually not decided in the same environment in which an RCT is run. Trial participation changes attention. Workflows are cleaner. Operators know they are being measured. Case mix, withdrawal technique, bowel preparation quality, and baseline ADR may not match the local service line that will use the tool at 4:30 p.m. on a Friday.

Controlled RCT setting contrasted with a busy real-world endoscopy suite

That distinction matters because pragmatic implementation studies have not consistently reproduced the RCT effect. Ladabaum and colleagues evaluated CADe in a U.S. integrated health system in a real-world workflow study of roughly 4,000 colonoscopies and did not find a statistically significant ADR improvement.[3] Levy and colleagues, as summarized in the same evidence landscape, also reported no statistically significant ADR gain after CADe deployment in normal practice.[2]

The CADILLAC trial makes the same problem harder to dismiss. As described by the National Cancer Institute, the trial enrolled 3,059 participants and did not show improvement in advanced adenoma detection, even though CADe trials more broadly have shown increases in overall adenoma detection.[4] For a purchasing committee, advanced adenoma detection is not a minor endpoint. It is closer to the clinical value story that justifies new hardware, staff training, device support, and contract spend.

This is where a relative-risk headline can mislead without being false. A system can improve ADR in a trial and still fail to move a health system’s local dashboard if baseline performance is already high, if alerts mostly identify lesions with lower near-term clinical consequence, if physicians override or habituate to alerts, or if the workflow change is too small compared with other determinants of colonoscopy quality.

A useful local evaluation therefore starts with the boring variables: baseline ADR by endoscopist, indication mix, withdrawal time practices if available, bowel preparation quality, pathology turnaround, procedure volume, surveillance capacity, and how the organization will define success. If the GI chair’s problem is a subset of endoscopists with persistently mediocre ADR, the trial literature may support a targeted intervention. If the service already performs well and is being asked to purchase CADe mainly because the category is visible, the pragmatic data become more important.

The Diminutive-Polyp Problem Changes the Value Question

Much of the CADe benefit in RCTs comes from adenomas smaller than 5 mm.[2] That does not make the benefit meaningless. Diminutive adenomas are still adenomas, and a technology that reliably prompts a second look may prevent some lesions from being ignored or passed too quickly. The issue is that “more adenomas” is not automatically the same institutional outcome as “better colorectal cancer prevention at acceptable downstream cost.”

AI highlighting tiny colon polyps beside symbols of increased surveillance burden and cost

The surveillance signal makes this concrete. A pooled analysis described in the evidence base found that the proportion of patients recommended for more intensive surveillance increased from 8.4% without CADe to 11.3% with CADe, corresponding to a relative risk of 1.35.[2] In a spreadsheet, that may look like a modest percentage-point change. In an endoscopy unit, it becomes repeat procedures, pathology volume, patient calls, scheduling slots, insurance exposure, and anxiety for people who are now labeled as needing closer follow-up.

Cost-effectiveness models are sensitive to exactly this point. Areia and colleagues estimated large potential U.S. annual savings under favorable assumptions, but the threshold conditions are the useful part for decision-makers: CADe had to raise ADR from 26% to at least 30%, or cost less than $579 per colonoscopy, to remain cost-effective in that modeling framework. A Spanish modeling study by Bustamante-Balén and colleagues reached a similar kind of conditional conclusion rather than a blanket economic endorsement.[5]

Those thresholds turn CADe from a general AI upgrade into a local math problem. If the incremental price is low, baseline ADR is suboptimal, and surveillance capacity is not already strained, the model may be plausible. If the per-procedure cost is high, baseline ADR is already strong, and the additional yield is concentrated in diminutive lesions that intensify surveillance, the same purchase can look very different.

This is also why product demonstrations can be unintentionally incomplete. The impressive moment is the green box around a subtle lesion. The missing half of the demonstration is the pathologist receiving more specimens, the scheduler finding follow-up capacity, the quality analyst deciding which metric moved, and the finance team discovering whether the modeled savings survive local pricing.

What an Institution Should Measure Before Calling the Added Yield a Win

  • Baseline ADR by individual endoscopist, not only the department average
  • Change in advanced adenoma detection, not only any-adenoma detection
  • Distribution of additional lesions by size, especially the share under 5 mm
  • Change in surveillance recommendations and follow-up colonoscopy demand
  • Incremental pathology workload and turnaround time
  • Actual per-procedure cost after device, maintenance, training, and support are included

None of these measurements requires hostility to CADe. They are the measurements that decide whether the detected lesions translate into the kind of value the organization thought it was buying.

Miss-Rate Evidence Keeps CADe Clinically Tempting

The strongest reason not to dismiss CADe is the tandem-study evidence. In tandem colonoscopy designs, the same patient effectively gives investigators a closer look at what was missed on the first pass. Wallace and colleagues reported that CADe reduced adenoma miss rates from about 31% to about 14%.[6] Glissen Brown and colleagues reported a reduction from about 37% to about 20%.[7]

Those findings speak to a different part of the adoption decision than a billing model or a departmental ADR report. They show that CADe can change visual recognition at the point where human performance is imperfect. For a quality committee, that is clinically hard to ignore. Even a careful endoscopist can miss a flat lesion, glance away at the wrong moment, or fail to register a subtle mucosal change while navigating a fold.

The tandem data do not erase the implementation problem. They explain why the software remains attractive despite it. A second look that lowers miss rates may be valuable even if the institution still has to prove that its own ADR, advanced adenoma detection, surveillance burden, and costs move in the right direction.

Deskilling Is Not Proven, but It Belongs in Governance

The deskilling concern is less mature than the RCT and tandem-study evidence, but it is not imaginary. Budzyń and colleagues reported an observational study in which endoscopists’ unassisted ADR dropped from 28.4% to 22.4% after they began routinely using AI, with p=0.009.[8] The study involved 19 endoscopists across 4 Polish centers, with a relatively low baseline ADR and no withdrawal-time data reported in the summary available for this appraisal.[8]

That design cannot establish a general causal law that CADe makes endoscopists worse. Geography, baseline performance, workflow, training culture, and unmeasured time factors all matter. Public coverage of the study appropriately made the concern visible, but visibility is not replication.[9]

Still, governance committees should not wait for a perfect deskilling literature before deciding what to monitor. If CADe becomes always-on infrastructure, the institution should know whether unaided performance changes over time. That may mean periodic CADe-off quality review, stratified ADR tracking, review of withdrawal technique, or simulation-based competency checks for clinicians who trained in an AI-supported environment.

The right posture is proportionate skepticism. Deskilling should not dominate the adoption decision today, but it should be written into the monitoring plan before the contract is signed. The harms most likely to be missed are the ones that look like ordinary workflow until several quarters of performance data make them visible.

Appraisal Constraints That Should Stay on the Record

Several limits should stay explicit in any institutional memo. The evidence base has substantial Asian and European representation, so generalizability to U.S. screening populations is not automatic.[2] The AGA’s 2025 living guideline made no recommendation on CADe because certainty around long-term colorectal cancer outcomes remains very low. That is a different question from whether CADe improves ADR; it asks whether the intervention has proved the endpoint that ultimately matters most.

There are also source-access limits. Several relevant full texts in Gastroenterology, Clinical Gastroenterology and Hepatology, and related journals are paywalled, which can restrict independent verification for teams without institutional access. This appraisal also does not include a first-party FDA MAUDE search, so device-event surveillance should not be inferred from the published clinical evidence alone.

FDA clearance belongs in the procurement record, but it should not be used as a substitute for clinical-effectiveness review. Clearance tells the organization that a product met a regulatory pathway for marketing. It does not answer whether the hospital’s endoscopists, patients, pathology department, scheduling template, and budget will experience net benefit.

A Conditional Decision Posture

CADe is reasonable to evaluate where the local conditions make benefit plausible: lower or variable baseline ADR, a credible plan to support less experienced endoscopists, acceptable per-procedure cost, available surveillance capacity, and agreement on which outcomes will define success. It is harder to justify as universal deployment when baseline performance is already strong, the business case depends on optimistic savings assumptions, or the organization has no plan to measure advanced adenoma detection, surveillance intensification, and unaided skill over time.

The current evidence supports a monitored implementation decision more than a settled standard-of-care declaration. Randomized efficacy, pragmatic effectiveness, lesion significance, downstream burden, cost thresholds, and deskilling surveillance all have to be held in the same frame. If any one of those streams is treated as the whole answer, the institution is not evaluating AI-assisted colonoscopy; it is buying the part of the evidence that is easiest to present.

References

  1. FDA Grants De Novo Clearance for First AI-Based CADe System for Colorectal Polyps, Applied Radiology
  2. Computer-aided detection and diagnosis systems for colorectal polyps: a systematic review and meta-analysis, Rabba et al., 2025
  3. Computer-Aided Detection of Polyps Does Not Improve Colonoscopist Performance in a Pragmatic Implementation Trial, Gastroenterology, 2023
  4. Artificial Intelligence for Colonoscopy: Promise and Pitfalls, National Cancer Institute, 2023
  5. Cost-effectiveness analysis of artificial intelligence-assisted colonoscopy in a colorectal cancer screening programme, Bustamante-Balén et al., 2025
  6. Impact of Artificial Intelligence on Miss Rate of Colorectal Neoplasia, Gastroenterology, 2022
  7. Adenoma Detection and Miss Rates During Colonoscopy With Artificial Intelligence, Clinical Gastroenterology and Hepatology, 2022
  8. Evidence-Based GI commentary on Budzyń et al. AI colonoscopy deskilling study, American College of Gastroenterology, 2025
  9. AI May Be Making Doctors Worse at Detecting Cancer, TIME, 2025

Risk-of-bias scorecard

Study design
RCTs, meta-analyses, pragmatic trials
External / prospective validation
Yes, pragmatic implementation studies
Key performance metric
Adenoma detection rate (ADR)
Overall rating
Moderate

Informational only — read the full disclaimer. This content supports procurement and research judgment, not clinical care decisions.

Submit a correction or sourcing issue

Blogarama - Blog Directory