A medical examiner office considering an “AI autopsy” product should ask a dull but decisive question before looking at the demo screen: what evidence for AI in forensic pathology and autopsy actually survives contact with medicolegal practice? The answer, on the current record, is narrow. The peer-reviewed literature supports controlled research and possibly a second-reader role under forensic pathologist oversight. It does not support routine deployment, replacement of the pathologist, or courtroom-ready reliance on an opaque model.
| Procurement question | Current evidence answer |
|---|---|
| Is AI in daily medicolegal forensic pathology practice? | No reviewed application was in daily medicolegal practice in the 35-application scoping review; all were rated at adapted technology-readiness level 2, meaning proof-of-concept rather than operational use. [1] |
| Are the headline performance numbers enough? | No. Reported accuracies can look strong, but the evidence base is small, unevenly validated, and weak on explainability. One 2025 systematic review found only 23% of analyzed studies exceeded 1,000 samples, and fewer than 15% used explainable-AI methods validated for legal contexts. [2] |
| Is there a forensic-pathology or autopsy AI tool with FDA clearance? | As of August 1, 2026, no FDA-cleared forensic-pathology- or autopsy-specific AI tool was identified. That regulatory finding is separate from the evidence verdict. |
| What use is currently defensible? | A constrained second-opinion, quality-assurance, or research role in which the forensic pathologist remains responsible for the report, the limitations, and the testimony. |

This is not an argument that the technology is useless. A model that makes a tired reviewer pause over a wound pattern, compare postmortem images more consistently, or surface an outlier in an immunohistochemistry workflow may have value. The problem begins when that narrow assistive possibility is sold as validated medicolegal readiness.
For related but narrower coverage of accuracy claims, see AI in cause-of-death medical investigations. The present appraisal is broader: it asks whether forensic pathology and autopsy AI, across applications, has the validation, calibration, replication, explainability, and regulatory footing needed for practice change.
The reviews are more important than the best-looking model
The most useful evidence here is not the single paper with the highest accuracy. It is the field-level view. Tournois and colleagues reviewed 35 AI applications in forensic medicine and found that none was used in daily medicolegal practice. Every application was assigned adapted technology-readiness level 2, placing the work in proof-of-concept territory rather than deployment. The same review found that 22 of 35 applications fell below its performance bar, a test set appeared in only about 58% of reports, and 27 of 35 reports had no quantified comparison with a non-AI gold standard. [1]
Those are not small footnotes. They change what the reported performance numbers mean. A model can classify a curated image set impressively and still fail the questions that matter to a forensic office: will it work on another jurisdiction’s cases, another scanner, another tissue-preparation workflow, another population, another injury mix, and another set of documentation habits?
Orsini and colleagues reached the same practical destination from a different review frame. Their 18-study PRISMA systematic review described a literature in which only 23% of analyzed studies exceeded 1,000 samples, explainable-AI methods validated for legal contexts appeared in fewer than 15% of forensic algorithms, and accuracy varied widely by task: about 67% for complex trauma, 94% for toxicological pattern analysis, 70–94% for deep-learning applications in neurological forensics, and 87.99–98% for gunshot-wound classification. [2]
That range is exactly why procurement discussions need discipline. “Up to 98%” is not the same claim as “externally validated, calibrated, independently replicated, explainable, and ready to affect a death investigation.” The former may be a promising research signal. The latter is a governance claim.

Application by application, the same evidentiary pattern keeps appearing
The field is not one thing. It includes postmortem imaging, wound classification, histology, immunohistochemistry, postmortem interval estimation, omics, toxicology, and image-based autopsy support. But the appraisal does not change much across categories: plausible signal, limited validation, sparse calibration, weak explainability, and no demonstrated routine medicolegal integration.
Postmortem imaging and autopsy images
Postmortem CT and autopsy-image classification are attractive targets because they look like familiar computer-vision problems. The model can search for patterns, compare pixels, and rank likely findings without fatigue. That makes sense as an engineering premise. It does not, by itself, resolve the forensic premise.
A death investigation is not just image recognition. The same scan or photograph may sit inside a chain of findings: scene information, medical history, decomposition, resuscitation artifacts, toxicology, histology, and competing manners of death. A model that performs on a selected imaging task still has to prove it can generalize across acquisition protocols, case mix, decomposition states, and institutional documentation practices. The review-level finding remains that published applications had not crossed into routine medicolegal use. [1][2]
Wound classification
Wound classification is where the temptation is strongest, because the reported accuracy range can sound procurement-ready. Orsini’s review reported deep-learning performance of 87.99–98% in gunshot-wound classification. [2] On a slide, that looks like an answer. In a report, it is only the beginning of a longer cross-examination.
The useful version of this tool is easy to imagine: the system flags cases where the wound class assigned by the human reviewer is discordant with image features seen in training data. That might justify a second look. It might also help with quality assurance in a large office where reviewers inherit variable image quality and incomplete histories. But a model trained on curated images is not automatically competent on blurred, blood-obscured, decomposed, surgically altered, or poorly lit evidence photographs.
The difference matters because wound interpretation can travel. It can affect manner-of-death classification, investigative direction, testimony, and a family’s understanding of what happened. If the model cannot explain which features drove its classification, and if its error profile has not been tested outside the original setting, the forensic pathologist still owns the conclusion without receiving much defensible support from the software.
Histology and immunohistochemistry
Histology and immunohistochemistry offer a more modest, and therefore more plausible, near-term use case. Pattern recognition, scoring assistance, and consistency checks are areas where digital pathology has obvious appeal. In a forensic IHC study, AI-assisted assessment reached 81.3% overall agreement with human scoring. [3]
That is a signal worth following, not a practice-changing endpoint. Agreement with human scoring is useful only after the reader knows which stains, tissues, artifacts, thresholds, and failure cases were involved, and whether the system was tested where it was not developed. In forensic pathology, the final question is not whether the tool agrees often enough to be interesting. It is whether disagreement is detectable, explainable, and safe under the pressure of a medicolegal report.
Postmortem interval estimation
Postmortem interval estimation deserves stricter treatment because the methodological warning is unusually direct. Sirago and colleagues appraised AI models for short-term PMI estimation of 72 hours or less using PROBAST and TRIPOD+AI-style criteria. The review identified only eight human studies, with sample sizes from 10 to 201, all small and single-center. It found high or unclear analysis risk of bias, no calibration reporting, and common unit-of-analysis problems. [4]
Calibration is not a statistical luxury here. If a model estimates PMI with impressive average error but is poorly calibrated at the time ranges that matter most to an investigation, the apparent precision can mislead. A stated interval can anchor a timeline, exclude or include witnesses, and shape investigative attention. Sirago’s review found that vitreous potassium and hypoxanthine models gave the most accurate continuous estimates, with error under three hours, but that finding still sits inside a small, single-center, calibration-poor literature. [4]
The safest current interpretation is not that AI has solved PMI. It is that a few biochemical signals may support research-grade continuous estimation under constrained conditions, while the evidence remains too fragile for unguarded operational confidence.
Omics, toxicology, and broader inference
Omics and toxicological pattern analysis are intellectually attractive because they promise inference from complex, high-dimensional signals. Orsini’s review reported accuracy as high as 94% for toxicological pattern analysis. [2] But high-dimensional data can also make overfitting look like discovery, especially when sample sizes are modest and external test sets are thin.
Forensic use also adds a chain-of-explanation requirement. A clinical or research model may be tolerated as a screening aid when its output is later checked by conventional methods. A medicolegal inference about a death investigation has to be interpretable enough for review, challenge, and testimony. The more abstract the input space becomes, the more the burden shifts to validation design, calibration, and explainability.
The same caution applies to adjacent NLP and prediction systems used in death investigation workflows. For readers tracking that neighboring domain, the appraisal of NLP in overdose-death investigation is a useful comparison: the output may help triage or quality review, but the evidence has to be separated from the operational claim.
FDA clearance is a separate question, and it does not rescue the evidence gap
Regulatory status and evidence quality should be kept in different boxes. A device can have clearance for one clinical use and still provide no support for forensic autopsy use. Conversely, a research model can have interesting evidence without being cleared for clinical or forensic deployment.
The useful contrast is PathAI’s AISight Dx. PathAI announced that AISight Dx received FDA 510(k) clearance K243391 on June 30, 2025, for primary diagnosis in clinical pathology. [5] That is a clinical digital-pathology clearance. It should not be read as clearance for forensic pathology, autopsy interpretation, wound classification, PMI estimation, or courtroom use.
As of August 1, 2026, no FDA-cleared forensic-pathology- or autopsy-specific AI tool was identified. Because this is a date-sensitive absence finding, governance teams should re-check clearance status at the time of procurement review and keep it separate from the literature appraisal. The site’s FDA clearance tracker is the natural place to monitor that distinction over time.
| Claim heard in review | What it would need to mean | What the current record supports |
|---|---|---|
| “FDA-cleared AI pathology platform” | Cleared for the specific intended use, population, workflow, and output being purchased | A clinical-pathology clearance exists for AISight Dx, but that does not establish forensic-autopsy readiness. [5] |
| “Validated on forensic data” | Externally tested on cases outside the development site, with documented error patterns | The anchor reviews describe a research-stage literature with limited external validation. [1][2] |
| “High accuracy” | Performance that generalizes and is calibrated for the decision being made | Headline accuracies vary by task and often come from small or proof-of-concept studies. [1][2] |
| “Explainable” | Output that a forensic pathologist can interrogate, document, and defend | Explainable-AI methods validated for legal contexts were rare in the reviewed forensic algorithms. [2] |
A courtroom problem, not just a workflow problem
In ordinary software procurement, a bad model may waste staff time or misroute work. In forensic pathology, a bad model can become part of a death investigation. It can affect how a family receives an explanation, how investigators frame a timeline, and how an expert witness is questioned.
That is why explainability cannot be treated as a cosmetic dashboard feature. If a wound model says “gunshot,” a PMI model says “within this interval,” or a histology model says “consistent with this scoring category,” the pathologist needs to know what features mattered, when the model fails, whether the case resembles the training population, and how uncertainty should be represented. A probability without calibration and context is not much help on the stand.
The signature on the report does not move to the vendor. Neither does the cross-examination.
Risk-of-bias scorecard for governance review
For a governance committee, the practical question is not whether a model is interesting. It is whether the evidence reduces enough uncertainty to justify a new operational dependency. On the current literature, the answer is usually no.
| Evidence domain | Current signal | Main weakness | Governance position |
|---|---|---|---|
| Field-level adoption | No reviewed AI application in daily medicolegal practice; all 35 applications in Tournois were at adapted technology-readiness level 2. [1] | Proof-of-concept status is repeatedly mistaken for operational readiness. | Do not procure as routine medicolegal infrastructure. |
| Study size | Only 23% of analyzed studies in Orsini exceeded 1,000 samples. [2] | Small datasets increase the risk that performance reflects local case mix, curation, or acquisition conditions. | Require external validation before deployment claims. |
| Test-set design | Tournois reported a test set in only about 58% of reports. [1] | Without credible held-out testing, performance estimates are unstable. | Treat internal accuracy as hypothesis-generating. |
| Comparator standard | Tournois found no quantified comparison with non-AI gold standards in 27 of 35 reports. [1] | A model cannot be judged operationally if the reference comparison is unclear or unquantified. | Require explicit comparator methods and disagreement analysis. |
| Calibration | Sirago found calibration was never reported in the short-term PMI studies it appraised. [4] | Uncalibrated probabilities or intervals can appear more precise than they are. | Do not use for timeline-sensitive operational decisions without calibration evidence. |
| Explainability | Orsini reported that fewer than 15% of forensic algorithms used explainable-AI methods validated for legal contexts. [2] | Opaque output is difficult to document, challenge, and defend. | Do not rely on black-box output as courtroom-ready evidence. |
| Regulatory status | No forensic-pathology- or autopsy-specific FDA-cleared AI tool was identified as of August 1, 2026. | Clinical pathology clearance should not be generalized to forensic autopsy use. | Track clearance separately from research performance. |
A vendor that wants to move beyond research support should be able to show, at minimum, external validation on independent forensic cases, calibration for the claimed output, prespecified failure analysis, human-comparator data, explainability appropriate to medicolegal review, and a regulatory position that matches the intended use. A strong retrospective accuracy number is not a substitute for that package.
The more defensible near-term path is controlled evaluation inside the office: silent-mode testing, second-reader experiments, discordance review, prospective auditing, and documentation of when the model changes a human conclusion. Even then, the model should be framed as an assistant that may draw attention to a case feature, not as the source of the medicolegal opinion.
AI in forensic pathology and autopsy is worth watching. It may become useful in narrow quality-assurance and second-opinion roles. But the burden of proof remains with anyone claiming routine deployment, replacement of the forensic pathologist, or courtroom-ready reliability.
References
- Artificial intelligence in forensic medicine: a scoping review, International Journal of Legal Medicine, 2023
- Artificial intelligence in forensic pathology: a systematic review, Frontiers in Medicine, 2025
- Artificial Intelligence Assisted Immunohistochemical Evaluation in Forensic Pathology, Diagnostics, 2025
- Artificial intelligence models for short-term postmortem interval estimation, Legal Medicine, 2026
- PathAI Receives FDA Clearance for AISight Dx Platform for Primary Diagnosis, PathAI, 2025