Skip to main content
ClinicalMind logoClinicalMind

The Evidence Behind AI for Drug Recall Detection

This appraisal examines whether NLP/ML systems applied to unstructured EHR text can reliably detect adverse drug events that precede recalls, reviewing the thin evidence base and the critical causality gap that limits current readiness for operational deployment.

Tool
Clinical NLP for ADE detection
Updated

Reviewer

Editorial Team

Clinical informatics editorial staff

FDA clearance status

Not applicable

A regulatory fact, reported separately from the evidence verdict.

Risk-of-bias verdict

High

The strongest argument for AI in drug recall detection does not start with a product dashboard or a claim about replacing safety reviewers. It starts with rofecoxib. Lependu and colleagues used NLP on clinical notes and reported that their system reproduced the rofecoxib signal, detected adverse drug events two years before official FDA alerts, and achieved an AUROC of 0.804 for drug-ADE associations.[1] That is the sort of result that makes the unused text inside EHRs hard to ignore.

It is also the sort of result that is easy to overread. A surfaced association is not a recall. A note-derived signal is not causality. And a single older study, however provocative, is not an operational benchmark for a pharmacovigilance program that has to decide what to escalate, what to adjudicate, and what to defend.

That tension is the useful starting point. NLP and machine learning applied to unstructured EHR text can widen the safety-signal net and may find clinically meaningful drug-event patterns earlier than structured sources alone. The current evidence, however, does not support buying or deploying these systems as standalone recall detection tools. They belong, at most, as governed surveillance inputs whose outputs still require human causality assessment and validation inside real safety workflows.

Clinical archive with illuminated text fragments rising from medical records toward a separated decision room

Why EHR Text Is Tempting

The temptation is not mysterious. A large share of clinically relevant information is not stored as clean, analyzable fields. The 2025 scoping review by Golder and colleagues notes that roughly 80% of EHR data is unstructured text, while these data remain rarely used in pharmacovigilance.[2] That means symptoms, temporal clues, medication changes, differential diagnoses, and clinician concern may sit in progress notes, discharge summaries, or consultation text while safety surveillance leans heavily on coded fields and spontaneous reports.

Spontaneous reporting systems have their own long-recognized under-reporting problem. Golder et al. cite evidence that only about 6% of adverse drug events are reported through spontaneous systems.[2] That figure does not prove that EHR-text NLP is ready for recall detection, but it does explain why safety teams keep looking for additional signal sources. If most adverse events never enter the conventional reporting stream, then a system that can scan routine care documentation has obvious appeal.

The practical value is not that NLP understands a chart the way a clinician does. It is that NLP can find candidate mentions at scale: a drug, a suspected event, a time relationship, a negation, or a note pattern that would be unrealistic to review manually across large populations. For post-market surveillance, that is a meaningful capability. It changes the first step from “wait for a report or a code” to “look for signals already embedded in care documentation.”

The Best Evidence Shows Signal Detection, Not Recall Detection

Lependu et al. remains the motivating case because it connects note-derived associations to a historically important safety issue. Their analysis of clinical notes identified drug-ADE associations and reproduced the rofecoxib signal, with ADEs detected two years before official FDA alerts.[1] In a field where delayed recognition has consequences, that finding deserves attention.

But the claim has to stay within its evidence boundary. The study supports the possibility that EHR text can surface earlier drug-event signals. It does not show that an NLP system can determine whether the drug caused the event, decide whether a label change or recall is warranted, or perform consistently across institutions, note styles, therapeutic areas, and surveillance teams.

That distinction matters because recall decisions are not triggered by mere co-occurrence. A patient may have joint pain after starting a drug because of the drug, because of the underlying disease, because of another medication, or because the symptom was already present and newly documented. EHR notes may contain the necessary clues, but extracting a drug and an event from a note is not the same as adjudicating the relationship between them.

What NLP/ML can plausibly contributeWhat remains unresolved for recall use
Extract drug and adverse-event mentions from unstructured notesDetermine whether the drug caused the event
Surface associations earlier than some existing channelsDecide whether an association meets regulatory action thresholds
Prioritize candidate charts or drug-event pairs for reviewReplace expert adjudication, epidemiologic assessment, or safety committee review
Improve surveillance sensitivity in settings where coded data are incompleteDemonstrate validated performance inside real recall decision workflows

The Comparative Evidence Is Promising, but Narrow

The most useful performance evidence is comparative, because it asks whether NLP adds value against an existing method rather than merely proving that a model can find something. Cai et al. provide that kind of comparison. In their vedolizumab-arthralgia analysis, NLP improved positive predictive value from 0.78 using ICD-9 codes to 0.90.[3]

That difference is operationally meaningful. False positives are not harmless in pharmacovigilance. They become chart-review queues, reviewer fatigue, safety committee agenda items, and sometimes expensive follow-up work. If NLP can reduce noise while retaining clinically important cases, it can make surveillance more usable rather than merely more sensitive.

Still, positive predictive value for one drug-event association is not a general recall detection result. PPV depends on the event definition, population, documentation practices, and prevalence of the target condition. A system that performs well for an arthralgia association in one setting may behave differently for hepatic injury, arrhythmia, severe skin reactions, or events that clinicians document inconsistently. The Cai comparison supports targeted utility; it does not supply procurement-grade generalizability.

Seven Comparative Studies Is Not a Mature Evidence Base

The field’s central weakness is not that the concept is implausible. It is that the comparative evidence base is thin. Golder et al.’s 2025 scoping review found only seven comparative studies meeting inclusion criteria for NLP and machine learning approaches to adverse drug event detection in EHRs.[2] That number is difficult to reconcile with confident claims that these systems are ready for broad operational deployment.

A small evidence base creates several problems at once. It limits confidence in external validity. It makes it harder to know whether performance is tied to a particular institution’s documentation style. It leaves uncertainty about which events are detectable, which drug classes are suitable, and how often signals would survive manual review. It also makes benchmark selection unusually important: a model that beats diagnosis codes in one comparison may still fall short when judged against adjudicated cases, longitudinal exposure data, or safety-review outcomes.

The absence of prospective validation inside recall decision-making workflows is especially important. Retrospective signal discovery can show that a system would have flagged a known concern earlier. It cannot show how a safety team would have handled the alert in real time, how many competing signals would have appeared the same week, how much review burden would have followed, or whether the signal would have been acted on appropriately.

The Causality Gap Is Not a Footnote

Golder et al. explicitly caution that NLP and machine learning systems cannot establish drug-event causality.[2] That is not a minor methodological caveat. It is the boundary between surveillance and regulatory action.

Infographic separating EHR note signal detection from human causality assessment

An NLP system can identify that a drug and an adverse event appear in a relevant textual relationship. It may even detect timing language, rule out some negated mentions, and rank associations that deserve review. But causality requires more: exposure timing, alternative explanations, background rates, dose relationships, dechallenge or rechallenge information when available, biological plausibility, confounding assessment, and expert judgment. Some of those elements may be partially available in EHR data. None becomes settled simply because a model scored an association highly.

This is where procurement language can become dangerous. “Recall detection” implies a readiness that “signal detection” does not. A recall is a regulatory and safety action built on evidence that a product presents a problem serious enough to remove or correct it. An EHR-text model may help identify where reviewers should look. It should not be treated as the entity that decides what has been proven.

What a Governed Pilot Would Need to Prove

A reasonable next step is not to dismiss NLP/ML for EHR-based drug safety surveillance. The rofecoxib finding is too important to ignore, and the structured-code comparison suggests real practical value in at least some use cases. The right question is narrower: what would a governed pilot need to demonstrate before an organization could rely on this output in safety operations?

  • A clearly defined target use case, such as screening for a specified drug-event pair or ranking candidate charts for reviewer triage.
  • A comparator that reflects current practice, not a weak straw man: diagnosis codes, spontaneous reports, manual review queues, or an existing pharmacovigilance process.
  • External validation across sites or populations where note style, coding behavior, and clinical workflows differ.
  • A human adjudication pathway that specifies who reviews model outputs, what evidence they examine, and how uncertainty is documented.
  • Prospective measurement of review burden, false positives, missed signals, escalation decisions, and downstream safety actions.

Those requirements are not bureaucratic decoration. They are the difference between an intriguing retrospective model and a safety process that can withstand clinical, legal, and regulatory scrutiny. A model that finds many candidate signals but overwhelms reviewers may reduce confidence rather than improve surveillance. A model that performs well at one academic center may not survive deployment in a different documentation culture. A model that cannot explain enough of its signal to support adjudication may create more uncertainty than action.

A Brief Cross-Context Check

Evidence from spontaneous reporting systems points in a similar direction, though it should not be merged casually with EHR-text NLP. Warner et al.’s 2025 systematic review found that ensemble machine learning methods can outperform traditional disproportionality measures in spontaneous reporting contexts.[4] That is relevant because it shows broader interest in machine learning for pharmacovigilance signal detection. It does not answer whether NLP on unstructured clinical notes is validated for recall workflows.

The distinction matters because spontaneous reports and EHR notes have different biases. Spontaneous systems depend on reporting behavior; EHR notes depend on documentation behavior. One may under-capture events that clinicians never report. The other may over-represent symptoms discussed during complex care, comorbid illness, or diagnostic uncertainty. Strong performance in one data environment does not automatically transfer to the other.

The Procurement Verdict

The evidence supports a careful yes for surveillance research and a firm no for standalone recall detection. NLP/ML applied to unstructured EHR text can surface drug-event associations that structured data and spontaneous reports may miss. It has produced a compelling early-signal example in rofecoxib and a useful head-to-head improvement over ICD-9 coding in a specific vedolizumab-arthralgia association.[1][3]

But the evidence base remains too small, too retrospective, and too removed from actual recall decision workflows to justify procurement as an autonomous or definitive recall detection capability. The most defensible role today is as an input to governed pharmacovigilance review: a way to find candidate signals earlier, focus chart review, and generate hypotheses for adjudication. The final safety judgment still belongs in a process that can assess causality, weigh uncertainty, and document why a signal does or does not merit escalation.

References

  1. Pharmacovigilance Using Clinical Notes, Clinical Pharmacology & Therapeutics, 2013.
  2. Leveraging NLP and ML for ADE Detection in EHRs, 2025.
  3. Cai et al. 2016 vedolizumab-arthralgia NLP versus ICD-9 comparison, 2016.
  4. Warner et al. 2025 systematic review of machine learning in spontaneous reporting systems, 2025.

Risk-of-bias scorecard

Study design
retrospective
External / prospective validation
No
Key performance metric
AUROC 0.804
Overall rating
High

Informational only — read the full disclaimer. This content supports procurement and research judgment, not clinical care decisions.

Submit a correction or sourcing issue

Blogarama - Blog Directory