Skip to main content
ClinicalMind logoClinicalMind

How reliable are AI diagnostic models for Lou Gehrig's disease?

The evidence behind AI diagnostic models for ALS is more uncertain than headline accuracy numbers suggest. This review of the 2024 meta-analysis explains why high pooled performance does not yet support clinical deployment.

Tool
AI diagnostic models for ALS
Updated

Reviewer

Editorial Team

Clinical informatics editorial team

FDA clearance status

None (no FDA-cleared ALS diagnostic AI tool)

A regulatory fact, reported separately from the evidence verdict.

Risk-of-bias verdict

High/Unclear

The best pooled estimate for AI in ALS research looks, at first glance, better than many clinicians would expect from a disease that still often takes too long to diagnose: 34 studies, pooled sensitivity of 94.3% and pooled specificity of 98.9% for AI-assisted ALS detection. In the same meta-analysis, gait-based models reported 91.2% sensitivity and 94.1% specificity, EMG-based models 92.6% sensitivity and 96.5% specificity, and MRI-based models 82.2% sensitivity and 77.3% specificity.[1]

Those numbers are encouraging enough to take seriously. They are not strong enough to treat as a deployment case. The current evidence appraisal is still research-stage: no FDA-cleared or FDA-approved AI tool exists specifically for ALS diagnosis, and the major systematic review found high or unclear risk of bias across much of the underlying literature.[1]

Abstract neural network brain separated from EMG and gait diagnostic equipment by a translucent divider

That distinction matters because ALS is exactly the kind of disorder that makes clinicians want better tools. The disease is rare, with about 30,000 patients in the United States, and the median diagnostic delay has been reported as 11 months.[2] During that interval, patients may move between primary care, orthopedics, spine imaging, neuromuscular referral, electrodiagnostic testing, and repeat examination before the pattern becomes clear enough to name.

A model that could shorten that interval would be valuable. A model that looks accurate only because it was trained and tested on convenient retrospective datasets could also misdirect referrals, reassure falsely, or label a patient with a devastating diagnosis before the evidence supports it. The difference is not semantic. It is the difference between a classifier and a clinical diagnostic tool.

What the 2024 meta-analysis actually supports

Umar et al. is the right place to begin because it is not a single-lab demonstration. It pooled 34 studies of AI models for ALS diagnosis and used QUADAS-2 and QUADAS-C to assess risk of bias and applicability.[1] That gives the review more weight than an isolated model-development paper, but it also makes the limitations harder to dismiss.

Evidence signalReported resultClinical interpretation
All included AI diagnostic studies34 studies; pooled sensitivity 94.3%; pooled specificity 98.9%Strong headline performance, but dependent on the design quality and populations of the included studies
Gait-based models91.2% sensitivity; 94.1% specificityPromising for movement-pattern analysis, but not equivalent to real-world diagnostic validation
EMG-based models92.6% sensitivity; 96.5% specificityClinically plausible because electrodiagnosis is already central to neuromuscular evaluation
MRI-based models82.2% sensitivity; 77.3% specificityLess convincing as a diagnostic signal in the pooled estimates
Regulatory statusNo FDA-cleared or FDA-approved AI tool specifically for ALS diagnosisNo current basis for independent clinical diagnostic deployment

Sensitivity and specificity in this setting answer a narrower question than procurement committees sometimes want answered. They describe how models classified cases and controls inside the studies that were available for pooling. They do not prove that the same models would perform safely across community neurology clinics, tertiary ALS centers, early presentations, mimic disorders, different EMG laboratories, different gait-capture systems, or populations that were not represented in training.

The risk-of-bias findings are therefore not a footnote to the pooled estimate. They are part of the estimate’s meaning. If patient selection is narrow, if controls are too clean, if cases are already obvious, or if the model is tested where it was developed, the task becomes easier than real diagnosis. A high score on that task may still identify a useful biological signal. It does not settle whether the model can help a clinician facing an uncertain patient.

This is especially important in ALS because the clinically relevant comparison is often not ALS versus healthy control. It is ALS versus cervical myelopathy, multifocal motor neuropathy, radiculopathy, motor-predominant neuropathy, inclusion body myositis, prior stroke, medication effects, functional weakness, or an early presentation that has not yet declared itself. A diagnostic model that has not been tested against those pressures may be learning a separation that is real but clinically incomplete.

Why external validation matters more than another decimal point

The most useful next study is not necessarily the one that reports 95% instead of 94% sensitivity. For ALS diagnosis, the more important question is transportability: whether the model keeps its performance when the institution, patient mix, acquisition hardware, feature extraction, referral timing, and comparator diagnoses change.

Rare disease research makes that difficult. Small ALS samples do not mean investigators are careless; they reflect the reality that ALS is uncommon. But rarity does not remove the need for external validation. It makes the validation problem more important, because the same limited cases can too easily shape model selection, feature engineering, and performance reporting.

The Umar review’s pooled performance should therefore be read as a signal that AI methods can extract ALS-relevant patterns from gait, EMG, imaging, and other clinical data, not as proof that a deployable diagnostic product exists.[1] That is a favorable research conclusion. It is also a restrained clinical conclusion.

Three-panel medical illustration of gait analysis, EMG waveform analysis, and MRI motor pathway analysis

The modalities are not equally persuasive

The modality breakdown is useful because it keeps the field from being treated as one undifferentiated AI claim. Gait models are attractive because ALS changes movement, balance, and motor control in ways sensors may quantify before a bedside impression becomes obvious. EMG and F-wave models are attractive for a different reason: they operate closer to a test neurologists already use when evaluating lower motor neuron dysfunction. MRI models, in the pooled analysis, looked less compelling as diagnostic classifiers, with lower sensitivity and specificity than gait or EMG approaches.[1]

That does not make gait or EMG models clinically ready. It does make them worth watching for different reasons. Gait approaches may be useful for triage, remote phenotyping, or longitudinal motor assessment if tested properly. EMG-based approaches may fit more naturally into existing neuromuscular workflows, but only if they improve decisions beyond expert electrodiagnostic interpretation and do so in the messy population that actually receives EMG.

MRI deserves a narrower reading. A model can find imaging patterns associated with ALS and still be a weak diagnostic instrument. The pooled MRI performance in the 2024 analysis does not justify treating imaging AI as a near-term standalone diagnostic route for Lou Gehrig’s disease.[1]

The strongest newer studies still do not close the deployment gap

Two recent studies show why the field should not be dismissed. They also show why the evidence still stops short of clinical deployment.

Martinez-Thompson et al. reported a Mayo Clinic F-wave AI model in Brain using data from 46,802 patients. The study found that a wavelet-based approach outperformed models based on clinical annotations.[3] That is a substantial sample for this domain, and the signal is clinically interesting because F-waves sit inside an electrodiagnostic context neurologists already understand.

But the design was retrospective.[3] A retrospective F-wave classifier can show that physiologic information was present in stored electrodiagnostic data. It cannot, by itself, show that clinicians using the model prospectively diagnose ALS earlier, refer more appropriately, avoid overcalling mimics, or improve patient-centered outcomes.

Zhao et al. approached the problem from a different direction: a blood-based gene-expression panel using XGBoost. The study reported 27- to 46-gene classifiers with 91% accuracy, along with survival prediction and identification of 8 repurposable drug targets.[4] That is closer to biomarker discovery than to a finished diagnostic service, and that distinction is not a criticism. Blood-based classifiers could eventually make ALS workups less dependent on late clinical pattern recognition. They could also help stratify patients for research.

Still, a gene-expression classifier with 91% accuracy is not the same as a validated diagnostic test ready for routine use.[4] Accuracy depends on the population tested, the comparator group, the timing of sampling, and how the model is calibrated for the clinical question. A biomarker model that separates known ALS cases from selected controls has not necessarily proved it can guide a first diagnostic decision in a patient with uncertain weakness.

Diagnosis, triage, monitoring, prognosis, and discovery are different claims

A recurring problem in AI-in-ALS discussions is that several legitimate use cases are collapsed into one word: diagnosis. They should be separated before any governance decision is made.

  • Diagnostic use asks whether a model can help determine whether a patient has ALS, usually in the presence of plausible alternatives.
  • Triage use asks whether a model can prioritize referral or further testing for patients who may need neuromuscular evaluation.
  • Monitoring use asks whether a model can track change over time after diagnosis or during follow-up.
  • Prognostic use asks whether a model can estimate survival, progression, or future clinical trajectory.
  • Biomarker-discovery use asks whether the model can identify biological patterns that deserve further validation.

Those uses do not carry the same evidentiary burden. A retrospective model may be quite useful for biomarker discovery. It may be reasonable to study as a triage aid if the downstream pathway includes specialist review and confirmatory testing. It is a much larger claim to say the same model can independently diagnose ALS.

The harm profile also changes. A monitoring model that adds noise to a research endpoint is a problem. A diagnostic model that falsely reassures a patient with early ALS may delay referral. A model that overcalls ALS may impose psychological and clinical consequences before the diagnosis has been properly established. Governance review has to evaluate the claimed use, not the most impressive result in the paper.

FDA status is separate from the evidence appraisal, but it points the same way

The regulatory point is simple: there is currently no FDA-cleared or FDA-approved AI tool specifically for ALS diagnosis. That does not prove the research is weak. FDA status and evidence quality are different questions. A model may have strong peer-reviewed evidence before clearance, and a cleared product may still require local validation and monitoring.

Here, however, the regulatory and evidence readings align. The field has promising retrospective classifiers, credible signals across several data types, and no prospective clinical utility study showing that deployment improves diagnostic decisions. For an AI governance committee, that is not a procurement-ready profile.

What would make the evidence clinically stronger

The next evidence threshold is not mysterious. ALS AI models need external validation in populations that resemble the intended use setting. They need comparator groups that include real diagnostic mimics, not only healthy controls. They need prespecified thresholds, transparent handling of uncertain cases, and reporting that makes false positives and false negatives visible rather than hiding them inside aggregate performance.

After that, the clinically meaningful test is prospective utility. Does the model shorten time to neuromuscular referral? Does it change the diagnostic plan? Does it improve accuracy when added to usual care? Does it increase unnecessary testing or overwhelm specialty clinics? Who reviews discordant results? What happens when the model is confident and the neurologist is not?

Those are harder studies than retrospective classification. They are also the studies needed before a health system should let an ALS diagnostic model influence clinical pathways at scale.

The current verdict

AI models for ALS diagnosis are credible research tools. Gait analysis, EMG and F-wave signal modeling, MRI classification, and blood-based gene-expression work all show that machine learning can detect disease-relevant structure in complex data. The 2024 pooled sensitivity and specificity estimates are too strong to ignore.[1]

They are also too weakly anchored to real-world validation to support independent clinical deployment. The evidence base remains limited by high or unclear risk of bias, small and selected samples, limited external validation, retrospective designs in leading newer studies, lack of prospective clinical utility evidence, and absence of an FDA-cleared ALS diagnostic AI tool.[1][3][4]

The clinical need is real: patients should not have to spend months in diagnostic uncertainty when earlier recognition is possible. The performance numbers are encouraging. But for now, AI models in ALS research support discovery, triage exploration, and future workflow studies more than deployment as diagnostic instruments for Lou Gehrig’s disease.

References

  1. Artificial intelligence in amyotrophic lateral sclerosis diagnosis: a systematic review and meta-analysis. PubMed. 2024.
  2. Diagnostic delay of amyotrophic lateral sclerosis. Scientific Reports. 2023.
  3. F-wave artificial intelligence model for amyotrophic lateral sclerosis. Brain. 2025.
  4. Blood-based gene expression panel using machine learning for amyotrophic lateral sclerosis. Nature Communications. 2025.

Risk-of-bias scorecard

Study design
Systematic review and meta-analysis
External / prospective validation
Limited; no independent prospective validation
Key performance metric
Pooled sensitivity 94.3%, specificity 98.9%
Overall rating
High/Unclear

Informational only — read the full disclaimer. This content supports procurement and research judgment, not clinical care decisions.

Submit a correction or sourcing issue

Blogarama - Blog Directory