Skip to main content
ClinicalMind logoClinicalMind

How Strong Is the Evidence Linking AI Outages to Patient Harm?

An evidence appraisal of what peer-reviewed studies actually prove about patient harm from AI clinical tool outages, distinguishing strong EHR-downtime evidence from thinner AI-specific data.

Tool
Multiple AI clinical tools
Updated

Reviewer

Editorial Team

Internal editorial review

FDA clearance status

Various

A regulatory fact, reported separately from the evidence verdict.

Risk-of-bias verdict

High

The short answer: plausible, important, not directly proven

The most defensible answer is narrow: peer-reviewed literature supports a plausible patient-safety risk when clinical systems become unavailable, but it does not yet contain a controlled study showing patient morbidity or mortality during an AI-specific clinical tool outage.

The strongest evidence comes from adjacent healthcare IT failures. Larsen et al. found that laboratory turnaround time increased 62% during electronic health record downtime, which is the cleanest kind of mechanism: the system is unavailable, a clinical process slows, and the delay can plausibly affect care decisions that depend on timely results.[1] A later study of the CrowdStrike outage found disruptions across 239 hospital services, again showing that digital infrastructure failure can interfere with care delivery at institutional scale.[2] The American Hospital Association’s 2024 Change Healthcare survey reported that 74% of hospitals experienced direct patient care impact from the cyberattack, although that survey does not prove patient injury or isolate AI tool unavailability.[3]

Evidence tierWhat it supportsWhat it does not prove
EHR and healthcare IT downtime evidenceClinical workflows can degrade when core digital systems are unavailable; delays and service disruption are measurable.It does not isolate AI clinical decision-support outages or measure AI-specific patient outcomes.
AI-enabled device recalls and adverse-event reportsAI-enabled medical technologies have documented recalls, malfunctions, injuries, and rare reported death events.Recall or adverse-event data are not the same as controlled evidence that downtime caused harm.
Model drift and prediction failure evidenceAn AI tool can be technically available while clinically unreliable.Prediction error is not the same as a declared outage, and these studies usually do not measure downstream morbidity or mortality.
Controlled AI-outage outcome studiesThis is the evidence governance teams would most want.In the sources reviewed here, this evidence is absent.
Three-tiered evidence framework showing EHR downtime evidence, AI device recall evidence, model drift evidence, and a question mark for missing controlled outcome studies

That distinction matters for procurement and governance. It is reasonable to plan as though an unavailable workflow-critical AI tool can affect patient care. It is not reasonable to cite the current literature as if it already proves that a specific AI outage caused excess mortality.

The strongest bridge is still EHR downtime

Larsen’s EHR downtime study deserves more weight than most dramatic outage anecdotes because it measures a workflow element that sits close to bedside decision-making. Laboratory turnaround time is not a soft perception measure. A delayed potassium, troponin, lactate, creatinine, or complete blood count can delay a treatment decision, a disposition decision, or an escalation decision. The study’s 62% increase in laboratory turnaround time during downtime does not, by itself, prove patient harm. It does show exactly where harm could enter the workflow: not through a vague “AI risk” category, but through a missing or delayed result.[1]

That is why EHR downtime is the best analog for AI outage planning. Many AI tools are now inserted into the same decision chain as the EHR, the lab interface, the imaging queue, or the alerting layer. If a stroke image flag fails to appear, if a deterioration score disappears from the rounding view, or if an oncology recommendation engine becomes unavailable during order entry, the question is not whether the outage has the label “AI.” The operational question is what step clinicians had been relying on, who notices its absence, and how long the replacement path takes.

The CrowdStrike evidence adds scale but less causal precision. A JAMA Network Open study found that the outage was associated with disruptions to 239 hospital services across affected institutions.[2] That is important for continuity planning because it shows how a nonclinical software failure can reach hospital operations. It is weaker than Larsen for patient-harm inference because “service disruption” is a broad endpoint. Some disrupted services may affect urgent care; others may affect administrative throughput or scheduled operations. Without patient-level outcomes, the number should not be inflated into a mortality claim.

The Change Healthcare cyberattack is similar. The AHA reported that 74% of surveyed hospitals experienced direct patient care impact.[3] That finding is valuable because payment, authorization, claims, eligibility, and pharmacy-related infrastructure can still alter care delivery even when they do not look like bedside systems. But it remains survey evidence from a major health system disruption, not a controlled measurement of AI downtime. It supports the practical premise that healthcare IT unavailability can affect care; it does not identify an AI model outage as the causal agent.

AI-specific evidence is closer to the technology, but farther from the outage question

The AI-specific safety literature should not be dismissed. It is the closest available body of evidence on how AI-enabled medical technologies fail after deployment. A 2026 JAMA Network Open analysis found that 43 of 903 FDA-cleared AI-enabled medical devices, or 4.8%, were recalled, with recalls occurring an average of 1.2 years after clearance.[4] That is a meaningful postmarket safety signal.

It is also not outage evidence. A device recall can involve labeling, software behavior, incorrect use, performance issues, or other safety concerns. Some recalled devices may be imaging algorithms or hardware-embedded systems rather than the clinical decision-support tools a hospital AI governance committee has in mind. A recall says a product had a recognized safety issue; it does not say the tool went dark during care, that clinicians lacked a fallback, or that a patient outcome worsened because of unavailability.

The adverse-event literature has the same tension. A JAMA Health Forum analysis reported 489 AI-related medical device adverse events, including 458 malfunctions, 30 injuries, and 1 death.[5] Those numbers should make governance committees uncomfortable enough to ask better questions. They should not be treated as a denominator-free proof that AI downtime kills patients. Adverse-event reporting is not the same as adjudicated causality, and the category includes AI-enabled medical devices broadly, not only AI tools embedded in clinical workflow.

The clinical validation gap is still relevant. The same JAMA Health Forum source reported that devices lacking sufficient clinical evidence at clearance had a hazard ratio of 3.33 for recall due to incorrect use.[5] That finding belongs in an AI procurement file because it links premarket evidence weakness to later recall risk. But the endpoint remains recall for incorrect use, not patient morbidity during an outage. The practical lesson is evidentiary humility: FDA clearance, clinical validation, safe integration, and downtime resilience are separate claims.

An AI tool can fail while still appearing available

Traditional downtime thinking starts with absence: the screen is unavailable, the interface is down, the order cannot be placed, the result cannot be seen. AI complicates that pattern because a model can continue returning outputs after its clinical usefulness has degraded. In that situation, the apparent system state is “up,” while the clinical function may be partially lost.

Censinet and GLACIS, writing from a third-party risk-management perspective, reported in July 2026 that 91% of AI models lose effectiveness over time, that failures can go undetected for 3 to 21 days, and that traditional IT monitoring catches only 15% to 20% of AI failures.[6] Those figures are useful for framing the continuity problem, but they should be read with the source in mind: this is vendor-published material from a company with commercial interest in AI risk management. It is not a neutral multicenter outcomes study.

The Epic Sepsis Model external validation is a better peer-reviewed example of model-performance failure, though still not an outage study. In a 2021 JAMA Internal Medicine external validation at Michigan Medicine, the model missed 67% of actual sepsis cases and had an 88% false-alarm rate.[7] That finding does not show what happens during a declared AI outage. It shows that a widely implemented model can perform poorly in a real deployment context, which is a different but related governance problem.

For downtime planning, this expands the definition of continuity. A hospital does not only need to know whether the AI endpoint is responding. It also needs to know whether the model is still clinically fit for the population, whether silent degradation would be detected, and whether clinicians understand when to revert to non-AI workflows. That is not evidence of injury. It is evidence that “available” and “safe to depend on” are not identical states.

What would stronger evidence look like?

A strong AI-outage patient-harm study would need more than a vendor incident notice and more than an adverse-event report. It would identify a specific AI clinical tool, define the outage window, describe what function disappeared, establish who was exposed, and compare clinical outcomes against an appropriate baseline or matched control period. It would also need to account for concurrent disruption. During a major cyber event, the AI tool may be unavailable at the same time as the EHR, lab interface, imaging system, paging system, pharmacy platform, or payment infrastructure. Attribution is not a clerical detail; it is the study.

The outcome should also match the tool. For an AI radiology triage system, the relevant measures may include time to image review, time to intervention, missed critical findings, or escalation delay. For a sepsis model, the measures may include time to antibiotics, fluid administration, ICU transfer, or mortality. For a documentation or coding assistant, patient-harm pathways may be weaker or indirect unless the tool affects orders, clinical reasoning, discharge instructions, or follow-up. Lumping all AI tools together makes the evidence look larger while making the safety question less answerable.

The denominator is equally important. Forty-three recalls among 903 cleared devices is interpretable because the denominator is visible.[4] A list of adverse events is harder to interpret without knowing deployment volume, reporting rates, exposure time, and whether events were independently adjudicated.[5] The same caution applies to outage anecdotes. A serious near miss can justify local action, but it should not be generalized into a frequency estimate without the population at risk.

How governance teams should use the evidence now

The absence of a controlled AI-outage outcome study is not a reason to leave AI out of downtime planning. It is a reason to be precise. If a tool is advisory, rarely used, and outside urgent workflows, its unavailability may be a nuisance. If it changes triage priority, suppresses or generates alerts, routes images, supports diagnosis, recommends treatment, or sits inside order entry, its unavailability belongs in patient-safety review.

  • Map the clinical function, not just the vendor product. Identify whether the AI tool affects triage, diagnosis, ordering, monitoring, discharge, follow-up, or billing-dependent access.
  • Define the fallback workflow before failure. A downtime binder that says “use clinical judgment” is not enough when the usual pathway includes an automated flag, score, queue, or recommendation.
  • Require outage and degradation notification terms. Contracts should distinguish complete unavailability, delayed output, stale data, model-performance degradation, and silent failure.
  • Preserve logs for reconstruction. After an incident, informatics staff need to know which recommendations were delivered, suppressed, delayed, ignored, or unavailable.
  • Tie risk tiering to workflow dependence. A low-risk model can become operationally high risk if clinicians have stopped performing the manual step it replaced.

Contracting language should also avoid the comfortable emptiness of “no known patient harm.” That phrase often means only that no harm has been confirmed or reported to the vendor. It may not mean a hospital reviewed delayed decisions, near misses, transferred workarounds, incomplete logs, or harm that emerged after the outage window. At the same time, governance documents should not convert every malfunction, recall, or model miss into a proven injury. Both habits make the evidence worse.

A defensible formulation

The best-supported statement is this: healthcare IT outages can measurably disrupt care processes, and AI-enabled medical technologies have documented safety failures, recalls, and adverse-event reports. Those facts make patient harm from AI clinical tool outages plausible, especially when the tool is embedded in urgent or high-dependence workflows. But the peer-reviewed literature reviewed here does not yet directly prove patient morbidity or mortality from an AI-specific tool outage.

That is enough to justify downtime planning, monitoring, fallback workflows, and stronger vendor obligations. It is not enough to claim that current studies have already measured the patient-outcome effect of AI outages. Plan as though AI unavailability can affect care when the tool is workflow-critical; cite the literature as analog, proximate, and functional evidence—not as a controlled AI-outage outcomes evidence base.

References

  1. Continuing Patient Care during Electronic Health Record Downtime — PMC, 2019.
  2. Patient Care Technology Disruptions Associated With the CrowdStrike Outage — JAMA Network Open, 2025.
  3. Change Healthcare Cyberattack Underscores Urgent Need to Strengthen Cyber Preparedness — American Hospital Association, 2024.
  4. Clinical Evidence and FDA Recalls of AI-Enabled Medical Devices — JAMA Network Open, 2026.
  5. Early Recalls and Clinical Validation Gaps in Artificial Intelligence-Enabled Medical Devices — JAMA Health Forum, 2025.
  6. The Connection Between AI Risk, Clinical Continuity, and Patient Harm — Censinet, 2026.
  7. External Validation of a Widely Implemented Sepsis Prediction Model — JAMA Internal Medicine, 2021.

Risk-of-bias scorecard

Study design
Observational
External / prospective validation
No
Key performance metric
Not reported in the cited evidence
Overall rating
High

Informational only — read the full disclaimer. This content supports procurement and research judgment, not clinical care decisions.

Submit a correction or sourcing issue

Blogarama - Blog Directory