For a procurement committee asking about evidence for AI in epilepsy diagnosis, the short answer is narrower than most product language implies. Deep learning models for detecting interictal epileptiform discharges on scalp EEG have reached near-expert performance in curated research datasets, and one major study showed a clinically meaningful specificity advantage over human consensus. That supports evaluation as a reader-assistive screening or triage tool. It does not support treating the model output as an independent epilepsy diagnosis.
That distinction matters before any accuracy number is discussed. A seizure-alerting wearable, a research IED detector, and a diagnostic EEG interpretation system are not the same regulatory or clinical object. In the evidence reviewed here, no FDA-cleared diagnostic AI tool for epilepsy was identified; the FDA-cleared epilepsy-adjacent devices commonly discussed in practice belong to seizure alerting rather than automated diagnosis from scalp EEG. Research systems such as SCORE-AI, IEDnet, and SpikeNet2 therefore need to be judged as clinical decision-support candidates, not as cleared diagnostic authorities.

The strongest affirmative case is specificity, not autonomy
SCORE-AI is the study that deserves to sit at the center of the discussion because it tests a question that actually matters in an EEG reading room: can an algorithm reduce overcalling without missing an unacceptable number of true epileptiform discharges? In the JAMA Neurology 2023 study, the model was tested across multiple centers and achieved an AUC range of 0.89 to 0.96 for automated interpretation of clinical EEGs.[1]
The more interesting result was specificity. SCORE-AI reached 90% specificity, compared with 73.3% specificity for three-expert consensus, while sensitivity was comparable rather than superior: 86.7% for AI versus 93.3% for human consensus.[1] In practical terms, the affirmative case is not that the model replaces the electroencephalographer. It is that, in the studied setting, the model was less prone to false-positive labeling.
That is not a minor advantage. False positives are not just a statistical nuisance in EEG. They create downstream review work, can push a chart toward an epilepsy label, and may trigger additional testing or treatment discussions that would not have existed if the tracing had been left as nonspecific or normal. A model that can reduce false-positive IED calls while maintaining near-expert sensitivity has workflow value, particularly in services where expert EEG review is a scarce resource.
The same numbers also define the boundary. Comparable sensitivity is not diagnostic independence, and higher specificity on a retrospective curated dataset is not proof that a model will perform safely when the patient is ventilated, sedated, neonatal, movement-contaminated, or represented by a technically poor study. SCORE-AI supplies a serious reason to evaluate AI-assisted screening. It does not settle deployment as a standalone diagnostic system.
| Evidence question | What the cited evidence supports | What it does not establish |
|---|---|---|
| Can deep learning detect IEDs at near-expert levels in curated scalp EEG datasets? | Yes, especially in SCORE-AI, with multicenter test performance and AUC 0.89-0.96. | It does not establish safe independent diagnosis across all clinical EEG populations. |
| Is there a clinically meaningful advantage? | SCORE-AI showed higher specificity than three-expert consensus, 90% versus 73.3%. | It did not show superior sensitivity to human consensus. |
| Can the evidence be generalized to difficult populations? | Only partially; important populations and EEG conditions remain underrepresented or excluded. | It does not close the gap for neonates, critically ill patients, or difficult artifact-heavy recordings. |
Where performance begins to fray
The cleanest limitation is not philosophical; it is anatomical and technical. Chung et al. reported that binary IED detection accuracy varied by brain region: 86.4% for frontal IEDs, 94.2% for temporal IEDs, and 97.2% for occipital IEDs.[2] The frontal decrement is exactly the kind of result that should slow a deployment committee down, because frontal EEG is also where eye movements and related artifacts can crowd the signal.

The authors attributed the weaker frontal performance plausibly to eye-related artifacts that the studied 2D CNN approach did not suppress well.[2] That is not a small edge case. In routine EEG review, frontal sharp transients, blink artifacts, and ambiguous anterior waveforms are precisely where an automated mark can become either useful triage or a false-positive burden.
This is also where broad claims about “AI accuracy” become too blunt. A global AUC does not tell the reader whether the model is reliable on a noisy frontal discharge, whether it behaves similarly in sleep and wakefulness, or whether it can keep its false-positive rate under control when the recording is full of ICU artifact. The evidence for frontal IED underperformance does not invalidate deep learning; it identifies the part of the reading queue where model output deserves more skepticism.
The same boundary applies to population coverage. SCORE-AI excluded neonates and critically ill patients, and those are not merely demographic footnotes.[1] Neonatal EEG has different maturational features, different background expectations, and different error consequences. ICU EEG adds sedation, metabolic encephalopathy, devices, movement, ventilators, and time pressure. A system trained and tested mainly outside those populations should not quietly inherit authority inside them.
The validation literature is impressive, but thinner than the headline AUCs suggest
The 2025 systematic review by Al-Breiki et al. widens the lens beyond a single strong model. It included 36 studies of AI-based epileptic spike detection, found that CNNs were the most common architecture at 55.6%, and reported peak AUC values as high as 0.99.[3] Those figures explain why the field looks mature from a distance.
The closer view is less procurement-ready. The same review found that 75% of studies relied on k-fold cross-validation without independent external validation, and small sample sizes were the most frequently cited limitation.[3] K-fold validation can be useful during model development, but it is not the same test as sending the model into a genuinely separate institution, recording environment, hardware mix, technologist workflow, patient population, and annotation culture.

That difference is especially important in EEG because the label is not a simple laboratory value. IED annotation depends on waveform morphology, field, state, montage, artifact context, and reader threshold. A model can perform well when training and test examples come from the same distribution yet lose reliability when the next site’s patient mix or annotation convention shifts.
Wang et al. provide broader context for AI in EEG analysis for epilepsy diagnosis and management, reflecting a field that is active across detection, classification, and clinical workflow support.[4] That breadth is encouraging, but it should not blur the narrower evidence question. Adoption of AI methods in epilepsy research is not the same as proof that automated IED detection improves patient outcomes when deployed prospectively.
What a defensible deployment would have to look like
The evidence supports a limited role: prioritizing studies, highlighting candidate IEDs, suppressing obvious false positives, or acting as a second-pass screen under human oversight. The model output should be visible as a suggestion with uncertainty, not as a final diagnosis. The EEG reader still has to decide whether the discharge has a convincing field, whether the morphology is epileptiform, and whether the finding is clinically relevant in the patient’s context.
A procurement review should therefore ask for evidence that resembles the intended use. If the proposed use is triage, the key measures include missed IED burden, review-time impact, false-positive marks per hour, reader override behavior, and whether urgent studies are surfaced faster. If the proposed use is diagnostic labeling, the evidentiary bar is much higher: prospective multicenter testing, predefined subgroup analysis, external validation in difficult populations, and evidence that clinical outcomes are not worsened by automation.
The most relevant subgroups are not hard to name. Neonates, critically ill patients, artifact-heavy studies, frontal discharges, rare epilepsy syndromes, and recordings from sites not represented in training all need explicit reporting. A pooled performance number that hides these groups is not enough for governance.
- Appropriate near-term use: reader-assistive screening, triage, candidate IED marking, and workload prioritization.
- Insufficiently supported use: autonomous epilepsy diagnosis or automated final EEG interpretation.
- Most persuasive current evidence: SCORE-AI’s multicenter performance and specificity advantage.
- Most important unresolved risks: subgroup degradation, artifact susceptibility, limited external validation, and absence of prospective outcome trials.
Procurement-grade verdict
The evidence for AI in epilepsy diagnosis is strongest when the claim is narrowed to automated IED detection support on scalp EEG. SCORE-AI shows that deep learning can approach expert sensitivity and may reduce false-positive overcalling in a way that matters operationally. Chung et al. show why site-specific and region-specific failure modes cannot be dismissed. The systematic review literature shows that many impressive AUCs still rest on validation designs that are weaker than clinical deployment requires.
A health system could reasonably evaluate these models as supervised screening or triage adjuncts, especially where expert EEG capacity is constrained and false-positive reduction is valuable. The current evidence does not justify standalone diagnostic deployment until prospective multicenter trials, independent external validation, and difficult-population performance reporting close the gap between curated-dataset success and the EEGs that actually arrive in the queue.
References
- Automated Interpretation of Clinical Electroencephalograms Using Artificial Intelligence. JAMA Neurology. 2023. https://pubmed.ncbi.nlm.nih.gov/37246210/
- Regional specificity of interictal epileptiform discharge detection using deep learning in electroencephalography. Scientific Reports. 2023. https://www.nature.com/articles/s41598-023-33906-5
- Artificial Intelligence for Epileptic Spike Detection in Electroencephalography: A Systematic Review. Journal of Epilepsy Research. 2025. https://pmc.ncbi.nlm.nih.gov/articles/PMC12185921/
- Artificial intelligence in EEG analysis for epilepsy diagnosis and management. Frontiers in Neurology. 2025. https://www.frontiersin.org/journals/neurology/articles/10.3389/fneur.2025.1615120/full