A cause-of-death AI claim can look impressive or disappointing depending on which number is placed on the slide: 62.65% final underlying-cause accuracy in one certificate-coding study, 0.978 test accuracy and 0.993 top-2 accuracy in another, 0.86 weighted AUC for structured EHR prediction in a preprint, or 74% immediate-cause identification by post-mortem CT against clinical information in a PMCT study.[1][2][3][4] Those figures are not a leaderboard. They are measurements from different operational jobs, against different reference standards, at different points in the mortality-data chain.
For evidence on ai in determining cause of death, the first question is therefore not “what is the accuracy?” It is “what exactly did the system determine, and who decided the answer was correct?” A model that assigns an ICD-10 chapter from a clean certificate line is not doing the same thing as a model that chooses the final underlying cause from a multi-condition causal chain. A verbal-autopsy classifier is not reading the same evidence as a post-mortem CT system. An EHR model predicting likely cause before the final death record exists is doing surveillance triage, not certifying death.

Four tasks sit behind one phrase
“AI determines cause of death” usually compresses at least four operational tasks. They share vocabulary, but not enough methodologic ground to support a single pooled accuracy claim.
| Task | What the model sees | Reference standard | What the headline metric actually measures | Evidence-appraisal implication |
|---|---|---|---|---|
| Death-certificate ICD-10 coding | Literal text and causal chains written on death certificates | Human-coded or nationally coded mortality records, sometimes with rule-based systems as comparators | Accuracy, top-2 accuracy, weighted precision/recall/F1, or sensitivity by disease group; performance changes sharply when the target moves from tentative cause to final underlying cause.[1][2][5][6] | Strongest evidence cluster, but the buyer must inspect task granularity and coder-in-the-loop design. |
| Verbal-autopsy classification | Structured questionnaire responses and sometimes free-text narratives from deaths without medical certification | Physician-coded or adjudicated verbal-autopsy causes | Cause-specific classification performance for surveillance categories, not legal determination of an individual death.[8][9] | Useful for population mortality surveillance; portability across settings is the central constraint. |
| Post-mortem imaging analysis | PMCT, CT-derived features, or forensic imaging data | Autopsy, forensic expert judgment, or imaging-autopsy concordance | Concordance or classification accuracy in small and heterogeneous forensic samples; PMCT and autopsy may disagree on legally important findings.[4][10][11] | Important but thinner evidence base; evidentiary defensibility matters as much as numeric performance. |
| EHR-based prediction | Structured EHR data and, in some models, clinical notes before or around death | Recorded death certificate or registry cause categories | AUC or classification performance for likely cause groups, often degraded when moved across institutions.[3] | A surveillance and case-finding instrument unless externally validated for the deployment population. |
This distinction is not academic hygiene. It decides who can safely use the output. A mortality coder needs a traceable code suggestion and a reason to trust it on rare chapters. A public-health analyst needs stable surveillance categories that do not collapse when transferred to another region. A forensic pathologist needs a defensible inference, not just a heat map over a small imaging set. A governance committee needs to know whether the reference standard was independent of the model output.
Certificate coding has the strongest evidence, and the widest apparent contradiction
The death-certificate coding literature is the most mature part of the evidence base because the task is operationally clear: turn reported death-certificate text into ICD-coded mortality data. It also produces the most misleading cross-study comparisons if the metric is stripped from the target.
Falissard and colleagues trained a deep neural network on approximately 8 million French death certificates and reported test accuracy of 0.978, compared with 0.745 for the Iris rule-based software; top-2 accuracy reached 0.993.[2] That is the kind of result that deserves attention: national-scale data, a large operational coding target, and a direct comparator. It is still not a universal “cause of death determined” result. It is a model evaluated in a specific certificate-coding environment against that environment’s coding standard.
Lee and Im’s Korean certificate study shows why the target definition can move the number more than the model label does. Their TRIPOD-AI-compliant KM-BERT model was trained and evaluated on 306,898 Korean death certificates. It reported 62.65% accuracy for final underlying cause, with an F1-macro of 0.1503, but 95.35% for tentative cause and 79.51% for individual-cause entry.[1] The same broad workflow therefore ranges from a modest final-answer result to a high intermediate-task result. Errors were concentrated in rare diseases and chapters requiring supplementary administrative information, which is exactly where a downstream coder would be least helped by a confident but poorly audited answer.[1]

Fang and colleagues add another practical warning. In 403,547 deaths from Fujian, a Wide and Deep model reported weighted precision of 95.75%, recall of 92.08%, and F1 performance that fell as the causal chain became more complex: 97.13% for single-cause chains and 79.50% for four-cause chains.[5] That degradation is not a footnote. It is the job. The certificate with a simple single-cause line is not the same coding problem as a certificate carrying multiple conditions, ordering questions, and underlying-cause selection rules.
AUTOCOD, evaluated in Portugal, is also useful evidence, especially because it reports disease-group sensitivity rather than relying only on a global accuracy figure. Ferreira and colleagues reported sensitivity of at least 0.90 for neoplasms, circulatory diseases, and respiratory diseases on 330,098 certificates, and performance held during excess-mortality periods.[6] The appraisal caveat is material: human coders had access to the model output, so the reference standard is not fully independent of the system being assessed.[6] For a fuller profile of that system, see the site’s AUTOCOD cause-of-death classification appraisal.
Certificate coding also exposes a separate quality-control use case: flagging unsuitable or illogical causal chains before they harden into surveillance data. The CODA challenge materials reported that 31.5% of Michigan 2022 resident deaths had NCHS-classified unsuitable underlying causes, and that 51.8% of AI prompts were invalid-causal-chain flags.[7] That does not prove an AI system can determine the true cause of death. It does show why coding-assistance tools may be valuable even when the final underlying-cause decision remains governed by coding rules, supplementary data, and human review.
The adjacent evidence in narrative death-investigation text is relevant but not identical. NLP can help classify overdose-related records, for example, without becoming a general mortality-coding engine; that boundary is covered in the site’s NLP overdose death investigation appraisal.
Verbal autopsy and EHR prediction are surveillance tools before they are adjudication tools
Verbal autopsy is often used where medically certified cause-of-death data are incomplete. The model is not checking a signed medical certificate; it is classifying cause categories from interview data. Idicula-Thomas and colleagues evaluated support vector machines on 18,826 childhood verbal autopsies and reported overall performance above 0.80, with 0.97 for diarrhoeal diseases.[8] The same source also noted that a randomized trial found physician coding outperformed most algorithms.[8] That combination supports a restrained conclusion: algorithms can help standardize population-level verbal-autopsy classification, but their performance is cause-specific and setting-dependent.
Free-text verbal-autopsy narratives add another layer. Jeblee and colleagues, summarized in a later diagnostics review, reported a narrative model precision of 0.773 and sensitivity/F1 of 0.77.[9] Those are not bad numbers for difficult narrative data. They are also not a substitute for knowing whether the deployment setting uses the same interview instrument, language, age distribution, and cause mix as the development data.
EHR-based cause-of-death prediction sits even earlier in the chain. Al-Garadi and colleagues reported, in a medRxiv preprint, structured-only weighted AUCs of 0.86 at Vanderbilt and 0.80 at Mass General Brigham, rising to 0.90–0.92 when clinical notes were added.[3] The same work reported sharp degradation on cross-institutional validation.[3] Because it is a preprint, and because portability is central to the claim, it should not be treated as settled evidence for procurement. It is better read as evidence that local EHR signals can be predictive, while transportability remains the hard test.
Forensic imaging matters, but the evidence is thinner
Post-mortem imaging is the most legally sensitive cluster in this comparison, and also the easiest to overstate. The reviewed imaging studies are not national mortality-coding systems. They are usually smaller forensic or radiology studies asking whether imaging, sometimes with AI assistance, can identify findings associated with cause of death or match autopsy conclusions.
Inai and colleagues reported that PMCT identified the immediate cause of death in 74% of cases, compared with 46% for clinical information.[4] That is an important result for post-mortem imaging workflows, but PMCT performance against clinical information is not the same claim as an AI system determining final cause of death across forensic contexts.
The broader forensic-AI literature remains feasibility-stage. Orsini and colleagues’ 2025 systematic review found that only 23% of included forensic AI studies had more than 1,000 samples, fewer than 15% used explainable-AI methods validated for legal contexts, and reported AI accuracy across forensic applications ranged from 67% to 94%.[10] The review also highlighted small imaging examples: a CNN for fatal cerebral hemorrhage with 0.94 accuracy in 81 cases, and a CNN for fatal head injury reporting 70% to 92.5% in 50 cases.[10] Those figures may justify continued research; they do not justify a generalized forensic cause-of-death claim.
The autopsy comparator is not a mere formality. Roberts and colleagues found a 30% major discrepancy between post-mortem imaging and autopsy in adult deaths, and all 10 pulmonary emboli were missed by imaging.[11] In a procurement review, that kind of miss pattern matters more than an aggregate accuracy percentage. A system that performs well on common visible injuries may still fail on a cause category with high legal and clinical significance.
Readers focused on autopsy and forensic imaging should use the site’s evidence gap for AI in forensic pathology and autopsy and cause-of-death medical investigation accuracy as the deeper cluster-specific appraisals.
A cross-cluster scorecard for governance review
A governance committee does not need a ceremonial average of incompatible metrics. It needs a risk-of-bias and applicability reading that keeps the model’s output attached to the study design. The pattern below follows the same practical logic as PROBAST+AI and TRIPOD-AI-style appraisal: population, predictor or input, outcome definition, reference standard, validation, transparency, and deployment fit.[12] The same evidence-appraisal posture is used in the site’s pancreatic cancer AI evidence and AI prostate cancer screening evidence appraisals.
| Appraisal domain | Certificate coding | Verbal autopsy | Post-mortem imaging | EHR prediction |
|---|---|---|---|---|
| Operational target | Usually clear: code literal text, assign tentative cause, or select underlying cause. | Clear for surveillance categories, less direct for individual adjudication. | Often varies by lesion, injury, or imaging-autopsy concordance target. | Predicts likely recorded cause categories from pre-existing clinical data. |
| Reference standard | National coding records or human coders; independence must be checked when coders see model output. | Physician-coded or adjudicated verbal autopsy; may vary across instruments and countries. | Autopsy or forensic expert judgment; discrepancies are clinically and legally meaningful. | Death certificate or registry labels; label quality depends on the downstream mortality record. |
| Validation maturity | Best developed, with large national or regional retrospective datasets and some held-out validation. | Moderate for selected settings; portability remains the main uncertainty. | Thin; many studies are small, heterogeneous, and not designed for legal explainability. | Promising local performance, but cross-institutional degradation is a major warning. |
| Most credible use | Coder assistance, coding quality control, surveillance production under defined workflow controls. | Population mortality surveillance where medical certification is limited. | Adjunctive forensic imaging interpretation, not stand-alone legal cause determination. | Surveillance triage, cohort finding, or early public-health signal detection. |
| Main procurement question | Which exact coding target was validated, and was the reference standard independent? | Was the model validated on the same population, language, questionnaire, and cause mix? | Can the output be explained, audited, and defended against autopsy-level discrepancies? | Has the model been externally validated on the institution where it will be used? |
This scorecard gives more credit to the certificate-coding literature than to the forensic-imaging literature because the former usually has larger datasets, clearer output definitions, and closer alignment with a production workflow. That is not a dismissal of imaging. It is a recognition that a 50-case or 81-case imaging model carries a different evidentiary burden from a national certificate-coding system.
It also keeps the 62.65% and 97.8% figures in their lanes. The lower Korean figure is attached to final underlying-cause selection, a harder and more administratively entangled target.[1] The higher French figure is attached to a different certificate-coding design and reference standard.[2] Reading one as failure and the other as proof of general cause-of-death determination would be the same error in opposite directions.
Regulatory status should not be inferred from production use
The reviewed materials did not identify an FDA-cleared medical device whose intended use is determining cause of death. That finding should be rechecked against the current FDA AI-enabled medical-device list and the site’s FDA Clearance & Incident Tracker before procurement or deployment, because regulatory status can change.
CDC MedCoder is a different category of fact. It appears in the HHS/ONC AI use-case inventory as HHS-CDC-00043, a system for coding literal-text cause-of-death information reported on death certificates, in operation since June 2022 within the National Vital Statistics System.[13] The inventory designation, including its “neither rights-impacting nor safety-impacting” classification, is not FDA clearance and should not be described as an independent safety determination.[13]
What the evidence can safely support
The evidence supports specific, bounded claims. AI can assist large-scale death-certificate coding when the target is defined, the coding standard is explicit, and performance is evaluated on data that resemble the deployment workflow. It can support verbal-autopsy surveillance where local validation confirms that the classifier travels across the relevant population, language, and cause mix. It can help explore EHR-based mortality surveillance, but current evidence is constrained by transportability and, for the cited EHR work, preprint status. It can contribute to forensic imaging research and possibly adjunctive workflows, but the reviewed imaging evidence is too small and heterogeneous to carry a broad cause-of-death determination claim.
The reviewed accuracy figures do not show that AI can determine cause of death across settings. They show that AI can perform narrower mortality-data tasks with very different ceilings, depending on whether the task is chapter-level coding, tentative-cause coding, final underlying-cause selection, verbal-autopsy classification, imaging-autopsy concordance, or EHR-based prediction.
References
- Development and Validation of a KM-BERT-Based Artificial Intelligence Model for Cause-of-Death Coding, Ewha Medical Journal, 2025.
- Performance of a Deep Learning Algorithm in the Coding of Causes of Death in French Death Certificates, JMIR Medical Informatics, 2020.
- Predicting Cause of Death From Electronic Health Records Using Machine Learning, medRxiv, 2025.
- Postmortem Computed Tomography for Determining Cause of Death, Virchows Archiv, 2016.
- Application of Artificial Intelligence in Automated Coding of Causes of Death — Fujian Province, China, China CDC Weekly, 2024.
- AUTOCOD: A Supervised Machine Learning System to Assist Mortality Coding in Portugal, JMIR AI, 2023.
- Using Generative AI to Validate Causal Chain of Death, Cal Poly Digital Transformation Hub.
- Machine Learning Methods for Automated Classification of Verbal Autopsy Data, BMC Public Health, 2021.
- Artificial Intelligence in Forensic Medicine and Forensic Sciences: A Systematic Review, Diagnostics, 2023.
- Artificial Intelligence in Forensic Medicine: A Systematic Review, Frontiers in Medicine, 2025.
- Post-mortem Imaging as an Alternative to Autopsy in the Diagnosis of Adult Deaths: A Validation Study, The Lancet, 2012.
- PROBAST+AI: A Tool to Assess Risk of Bias and Applicability of Prediction Model Studies Using Artificial Intelligence, BMJ, 2025.
- MedCoder: Coding Literal Text Cause of Death Information Reported on Death Certificates, HHS/ONC AI Use Case Inventory.