In litigation after a disease outbreak, AI usually does not arrive wearing a label that says “future evidence.” It arrives first as a dashboard that helped an epidemiology team prioritize interviews, a model that forecast hospital demand, an ambient scribe note from a fever-clinic visit, or a video clip offered to show what someone said during a public health response. The legal problem begins later, when a party asks a court to treat that output as proof.

That is the practical question in legal cases involving AI and disease outbreaks: not whether courts should admire or fear AI, but whether the material can be authenticated, tested, explained, preserved, and challenged. Existing evidence law already has tools for that inquiry. The harder part is that outbreak conditions often produce exactly the kind of recordkeeping gaps that evidence law punishes later.

Epidemiological AI data visualizations transitioning into courtroom evidence

When Outbreak Tools Become Exhibits

Outbreak-related AI evidence is not one category. A court may see AI-assisted surveillance outputs, machine-learning epidemiological models, AI-generated clinical documentation, automated summaries of public health records, or synthetic media. Those materials differ in purpose, provenance, and risk.

Material likely to appear in outbreak litigationWhy it matters in courtPrimary evidentiary pressure
AI-assisted aggregation of case reports, laboratory feeds, or syndromic surveillance signalsMay support timelines, notice, causation theories, or reasonableness of public health responseData lineage, human validation, completeness, and bias in reporting sources
Epidemiological AI models or outbreak forecastsMay be offered through an expert to explain spread, predictability, allocation choices, or alternative interventionsTestability, validation, error rates, peer review, model transparency, and fit to the legal question
AI-generated clinical notes or scribe documentationMay become part of the medical record in malpractice, negligence, coverage, or institutional response disputesAuthentication, preservation of source recordings, edit history, and whether deletion created discovery problems
AI-generated legal or factual submissionsMay contaminate pleadings, briefing, or expert materials with unsupported assertionsVerification, sanctions risk, and whether cited authorities or facts exist
Deepfake audio or videoMay be offered to prove statements, conduct, consent, warnings, or public communicationsAuthenticity, chain of custody, forensic examination, and witness corroboration

The lower-risk end of the spectrum is not risk-free. A dashboard that aggregates laboratory reports can still inherit delayed reporting, unequal access to testing, inconsistent coding, or missing demographic data. But if a public health team can show where the data came from, who reviewed the output, what corrections were made, and what uncertainty was preserved, the evidentiary fight looks familiar.

The higher-risk end looks different. A clinical note generated from an ambient recording may be the only contemporaneous account of a visit. A synthetic video may appear facially convincing. A legal filing may cite a nonexistent case with the confidence of a reporter citation. In those settings, the output can look cleaner than the process that produced it, and that is where courts should slow down.

The Existing Rules Are Usable, But They Need to Be Applied Carefully

The most useful starting point is not a new AI-specific panic rule. Grimm, Grossman, and Cormack’s framework for “Artificial Intelligence as Evidence” is valuable because it keeps the analysis inside recognizable evidence questions: relevance, authenticity, hearsay, original-writing concerns, expert reliability, prejudice, and the ability of the opposing party to challenge the material.[1]

That structure matters in outbreak litigation because AI evidence may play different roles. A model may be the basis for an expert opinion. A surveillance output may be a business or public record. An AI scribe note may be offered as part of a medical record. A synthetic video may be demonstrative, impeachment material, or purported direct evidence. The same word, “AI,” does not answer the threshold question: what is the exhibit being offered to prove?

For epidemiological AI models, the Daubert-style inquiry should do real work. A party relying on a model should be prepared to address whether the model can be tested, whether validation evidence exists, whether known or potential error rates have been measured, whether the method has been peer reviewed, and whether the relevant scientific community accepts the approach for the use being claimed. Naming those factors is not enough. The court needs to know what was tested, on what data, under what outbreak conditions, and against what benchmark.

Fit is often the neglected issue. A model might be useful for allocating staff during a fast-moving outbreak and still be a poor fit for proving legal causation in a negligence suit. A tool optimized to detect a cluster early may tolerate false positives that are reasonable for surveillance but problematic when used to assign responsibility. A forecast built to guide resource planning may not answer whether a particular defendant’s action caused a particular plaintiff’s exposure.

Explainability also matters, though it should not be reduced to a demand that every model be simple. Some complex models may be admissible if their inputs, validation, limitations, and outputs can be meaningfully challenged. The problem is not complexity by itself. The problem is an evidentiary black box that asks the court to accept a conclusion without enough information to cross-examine the path from outbreak data to courtroom claim.

Bias and Incomplete Surveillance Data Are Reliability Problems, Not Public Relations Problems

Outbreak datasets are rarely pristine. Testing access may vary by geography. Mild cases may be undercounted. Records may lag. Certain workplaces, institutions, or communities may appear more heavily represented because they were monitored more aggressively, not because they were necessarily the origin of spread. If an AI model learns from that uneven record, the legal question is not simply whether the software is advanced. It is whether the resulting output is reliable for the proposition offered.

That inquiry should be concrete. What data sources were included? What sources were excluded? Were missing values imputed, ignored, or flagged? Did humans review anomalies? Were uncertainty intervals preserved in the final report, or did the litigation exhibit convert a probabilistic output into a categorical statement? The answers can affect admissibility, weight, or both.

AI Scribe Notes Create a Preservation Problem Before Anyone Files Suit

Clinical documentation is where the evidentiary problem becomes less theoretical. AI scribes are designed to make care easier: record the encounter, generate a draft note, and let a clinician review or edit it. In an outbreak, that can be a practical relief for overwhelmed clinicians. Later, the same note may become evidence of symptoms, warnings, consent, triage decisions, infection-control instructions, or follow-up plans.

AI clinical transcription interface linked to legal discovery materials

Dr. Katherine E. Goodman’s work at the University of Maryland School of Medicine identifies a litigation-sensitive practice that deserves attention: healthcare systems deleting AI scribe recordings to avoid malpractice discovery risks.[2] That may be framed internally as data minimization. In litigation, however, deletion can become the fact everyone has to explain.

The tension is real. Health systems have legitimate reasons not to retain every raw recording indefinitely. Recordings may contain sensitive information, third-party statements, incidental disclosures, or material beyond what belongs in the legal medical record. But once litigation is reasonably anticipated, routine deletion practices can collide with preservation duties. In an outbreak setting, that anticipation point may arrive early: a cluster in a facility, a widely reported exposure event, a serious adverse outcome, or a formal complaint can all change the preservation analysis.

Authentication then becomes more than proving that a note sits in the electronic health record. A party may need to show how the note was generated, whether a recording existed, whether the clinician reviewed the AI draft, what edits were made, when they were made, and whether the final note accurately reflects the encounter. If the recording was deleted before those questions arose, the remaining audit trail may carry more weight than anyone expected when the tool was deployed.

This is not an argument that every AI scribe recording must be retained forever. It is an argument that retention cannot be treated as a purely technical setting. The records team, counsel, clinical leadership, and vendor should understand what evidence exists at each stage: raw audio, transcript, AI draft, clinician edits, metadata, deletion logs, and final note. If the system destroys the most probative layer before legal duties are evaluated, the health system may preserve the polished output while losing the material needed to authenticate it.

Hallucinated Authorities Show Why Verification Cannot Be Assumed

Courts are already seeing AI-generated falsehoods in legal submissions. A March 2026 Drug & Device Law Blog discussion described more than 350 documented cases involving self-represented litigants citing AI-hallucinated cases, plus more than 200 instances involving lawyers.[3] Those figures do not measure outbreak litigation specifically, and they should not be stretched beyond what they show. They do show that fabricated legal authority is no longer a rare classroom hypothetical.

AI-generated legal citations fragmenting beside courtroom documents

For outbreak cases, the lesson is broader than lawyer discipline. If a filing can contain nonexistent cases, an expert report can contain unsupported AI-generated summaries, and an administrative record can include automatically produced text that no one checked against source documents. Verification needs to attach to the specific use: legal citations, factual summaries, medical-record abstracts, or epidemiological conclusions.

The failure mode is familiar to anyone who has reviewed public health records under deadline. A generated summary may be mostly right and still omit the qualification that matters. It may merge dates, flatten uncertainty, or cite a document that says something narrower than the summary claims. In court, “mostly right” can still create an admissibility fight if the proponent cannot trace the assertion back to the underlying record.

Deepfakes Are a Proximate Warning, Even If Outbreak Cases Have Not Yet Turned on One

Synthetic media raises a different kind of authentication problem. Mendones v. Cushman & Wakefield, described in AI-generated evidence scholarship as one of the first instances where AI-generated deepfake video was submitted as authentic evidence in court, is not an outbreak case.[4] It is still a useful warning. Courts can no longer assume that apparent visual realism carries the authentication force it once did.

In outbreak litigation, deepfake scenarios remain largely hypothetical. A court might someday face a disputed video of a facility manager giving instructions, a public official making a statement, or a clinician discussing triage. The point is not to predict frequency. The point is to recognize that authentication will need more than a witness saying the video looks plausible. Chain of custody, device metadata, forensic review, source-system records, and corroborating witnesses may all become necessary to decide whether the exhibit is what its proponent claims.

Proposed Rule 707 Is a Signal, Not Current Law

Reform proposals are emerging because these disputes are no longer confined to academic panels. The National Center for State Courts has discussed Proposed Rule of Evidence 707, which would subject machine-generated evidence to the same reliability standards as expert testimony.[5] That proposal is important, but it is not operative law. Courts still have to work with the evidence rules they have.

The appeal of a rule like Proposed 707 is that it would force reliability questions to the surface when a machine output is offered as evidence. But even without it, litigants can ask many of the same questions through existing doctrine: relevance, authentication, hearsay exceptions, expert admissibility, Rule 403-style prejudice concerns, discovery rules, and sanctions principles when evidence has been lost or fabricated.

The risk is that a new rule becomes a shortcut label rather than an operational discipline. If a hospital cannot explain its AI scribe retention settings, or a health department cannot reconstruct the data sources behind a model output, the courtroom problem will not be solved by citing an AI evidence rule. The missing link will still be missing.

Deployment Choices Decide the Later Evidence Fight

Public health agencies and health systems make many of the decisive evidentiary choices before counsel drafts a complaint or subpoena. The Network for Public Health Law’s guidance on generative AI and health departments frames those choices as legal-risk questions at deployment, not merely technical procurement issues.[6]

For outbreak AI tools, that means asking evidence questions early. Who owns the source data? What vendor logs are retained? Can the agency export inputs, outputs, prompts, model versions, and human edits? Does the contract allow access during litigation? Are uncertainty measures preserved? Are staff trained to distinguish a machine-generated summary from a verified finding? The answers determine whether later testimony will be grounded in a record or in institutional memory.

A useful deployment file for an outbreak AI system would not need to be theatrical. It would identify the tool, its purpose, the data feeds used, validation steps, known limitations, human review points, retention settings, vendor responsibilities, and escalation triggers for legal holds. That file would not make the AI output automatically admissible. It would give the records custodian, epidemiologist, clinician, or risk manager something concrete to defend.

What Courts Should Ask When AI Outbreak Evidence Is Offered

A disciplined court inquiry does not need to begin with whether AI is good or bad. It can begin with ordinary evidence questions stated precisely.

  • What is the AI material being offered to prove: notice, causation, diagnosis, timing, damages, authenticity, or expert opinion?
  • Who generated or controlled the output: a public health agency, hospital, vendor, clinician, expert, attorney, or litigant?
  • Can the proponent trace the output to source data, model version, prompt or configuration, human review, and final exhibit?
  • Was the method tested for the use being claimed, and are validation results or error rates available?
  • Were material uncertainties, missing data, limitations, and possible biases preserved or stripped away?
  • Has any underlying recording, metadata, draft, audit log, or source file been deleted, altered, or placed outside the party’s practical control?

Those questions will not produce the same answer for every exhibit. An AI-assisted line list checked by trained epidemiologists may raise manageable foundation issues. A model used beyond its validated purpose may be excluded or narrowed. An AI scribe note without audit history may still be part of the medical record, but its weight and completeness may be contested. A deepfake allegation may require forensic evidence before the court can decide whether the jury should see the media at all.

Courts do not need to invent a separate evidentiary universe for AI-generated outbreak materials. They do need parties to bring more than polished outputs. In outbreak litigation, the reliable exhibit is not just the graph, note, forecast, or video. It is the traceable path behind it: the source data, the human checks, the retained metadata, the preserved uncertainty, and the record of what was deleted and why. Existing law can handle that path. Current operational habits are not yet consistently built to preserve it.

References

  1. Artificial Intelligence as Evidence, Northwestern Journal of Technology and Intellectual Property, Vol. 19, Iss. 1.
  2. UMD Faculty Profile, University of Maryland School of Medicine, 2026.
  3. AI Hallucinations in Court: A Case Study in How Bad It Can Get, Drug & Device Law Blog, March 2026.
  4. Mendones v. Cushman & Wakefield, case record, 2025.
  5. AI-generated evidence is a threat to public trust in the courts, National Center for State Courts.
  6. Generative AI and Health Departments: Legal Considerations and Risks, Network for Public Health Law.