Skip to main content
ClinicalMind logoClinicalMind

What the evidence shows about AI in MS drug discovery

Most AI in MS drug discovery remains hypothesis-generating; the strongest peer-reviewed evidence supports treatment-response prediction and trial-data enrichment, not molecule discovery. This appraisal separates defensible patient-level claims from hypothesis-generating discovery research, giving procurement and research teams an evidence-tiered judgment they can apply to vendor pitches.

Tool
AI in multiple sclerosis drug discovery
Updated

Reviewer

Editorial Team

Editorial Team

FDA clearance status

No FDA clearance reported; primarily research/development tools

A regulatory fact, reported separately from the evidence verdict.

Risk-of-bias verdict

Moderate-to-high

A procurement team hearing an “AI-discovered MS drug” pitch in Q3 2026 should ask for the pipeline stage before asking for the demo. The peer-reviewed evidence on AI in multiple sclerosis drug discovery does not support one flat claim. It supports a tiered one: AI is useful for generating targets, ranking repurposing candidates, stratifying patients, predicting treatment response, and enriching trials; in the sources reviewed here, it has not yet shown prospective clinical validation of an AI-nominated multiple sclerosis drug.

The patient-facing claims are strongest where the model is close to an actual treatment or trial decision. A 2025 Nature Communications study used machine learning on high-content imaging of CD8+ T cells to predict natalizumab response, reporting 92% accuracy in discovery and 88% in validation; the same paper notes that about 35% of natalizumab users do not respond completely within two years [1]. A 2022 Nature Communications trial-enrichment study analyzed six randomized controlled trials with 3,830 participants and found that an AI-guided predicted-responder subgroup carried a clearer anti-CD20 signal in a held-out primary progressive MS test set than the full group [2]. Those are not proof that AI discovered a drug. They are evidence that AI may help decide who is likely to benefit or how a trial population is analyzed.

The discovery-side literature is more preliminary. A PRISMA review of 70 AI-in-MS papers from 2018 through 2022 found only seven categorized as treatment selection or optimization [3]. That review is useful because it shows where the field’s published center of gravity was: diagnosis, imaging, classification, monitoring, and prediction more than validated drug selection.

Three-stage evidence-tiered pipeline for AI in drug discovery, from hypothesis-generating target work through patient-level evidence to clinical validation not yet reached

The evidence map depends on what the model is being asked to change

A model that ranks a protein target has a different evidentiary burden from a model that decides whether a patient enters a trial stratum or receives natalizumab. The first may be valuable with genetic triangulation, pathway coherence, and plausible druggability. The second needs discrimination, calibration, validation, and preferably prospective evidence that using the model improves decisions rather than merely predicting outcomes already visible in retrospect.

Pipeline positionWhat the reviewed evidence showsPatient-facing consequenceAppraisal judgment as of Q3 2026
Target prioritization and causal-protein discoveryMulti-omics, Mendelian randomization, colocalization, and protein-QTL studies nominate proteins and pathways for follow-up.Could influence which molecule or pathway gets tested next, but does not itself change treatment.Scientifically useful hypothesis generation; not clinical validation.
In-silico repurposing and compound rankingQSPR, ranking, and drug-repurposing screens produce candidate lists, including non-MS drugs or supplements in some studies.Could guide laboratory work or portfolio review; should not be presented as a ready MS therapy.Early discovery evidence; wet-lab and prospective clinical validation remain the missing steps.
Treatment-response predictionThe strongest reviewed example predicts natalizumab response from patient-derived CD8+ T-cell imaging with discovery and validation performance reported.Could eventually affect which patient receives natalizumab or another disease-modifying therapy.Closer to clinical action, but still needs prospective utility before routine use.
Trial enrichment and responder stratificationRetrospective trial analyses show AI can concentrate treatment effects in predicted-responder groups.Could change enrollment, power, subgroup interpretation, or trial design.More defensible than molecule-discovery claims; still not evidence that AI invented the therapy.
Disease reclassificationLarge-scale trial-data models reclassify MS into data-driven states and dimensions with external validation.Could reshape trial analysis, enrichment, and disease monitoring.Important adjacent evidence; not drug discovery.
Vendor partnershipsPublic collaborations name MS as a focus area but do not yet provide peer-reviewed MS outputs in the reviewed record.No patient or formulary consequence until data appear.Treat as pipeline activity, not evidence.

Discovery-stage papers are useful when they stay in their lane

The most tempting overstatement in this field begins with a legitimate discovery result. A target-prioritization paper finds a protein, a repurposing screen ranks a compound, or a multi-omics model triangulates a pathway. The press-ready version then becomes “AI found the next MS drug.” That sentence skips the work that matters most to patients: biological validation, dose and exposure, safety in the intended population, comparative effectiveness, and prospective trial evidence.

Jiang et al. is a good example of a study that is interesting precisely because it is careful about this boundary. The 2026 Journal of Neuroinflammation paper prioritized 48 proteins and 13 non-MS drugs, identified RRM2B as a shared target of cladribine and clofarabine, and reported seven nutrient or supplement candidates. It also explicitly framed the work as hypothesis-generating. The caution matters: only one of six key proteins reached Bonferroni significance, and the brain pQTL component was drawn from about 400 postmortem samples [4].

There is nothing trivial about that work. If a program team is deciding which proteins deserve assay development, animal-model exploration, or medicinal-chemistry attention, genetic and proteomic triangulation can be useful. But the output is a prioritized experiment list. It is not evidence that clofarabine, a supplement, or any other candidate should enter MS care because an AI or integrative model placed it high on a list.

Farooq et al. sits even earlier in the chain. The 2025 BMC Chemistry study used QSPR modeling with TOPSIS and WASPAS ranking across 12 MS drug candidates, but the reviewed record describes no wet-lab validation and notes small-dataset limitations [5]. Ranking 12 candidates can be a useful computational exercise. It does not tell an MS clinician which patient should switch therapy, nor does it tell a sponsor that the top-ranked compound has survived pharmacology, toxicity, or efficacy testing.

Lin et al. used Mendelian randomization and colocalization to report five proteins as promising MS targets [6]. That type of design is valuable because it can reduce some confounding that burdens observational biology. It still does not prove that modulating the protein with a drug will produce a safe, clinically meaningful benefit in relapsing or progressive MS. A genetically supported target is a stronger starting point than an unsupported one; it remains a starting point.

The 2025 PNAS multi-omics target-prioritization study belongs in the same evidence tier. Multi-omics integration can help a sponsor or academic group decide where to look next, especially when transcriptomic, genetic, proteomic, and disease-association signals converge [7]. The value is in narrowing the search space. The missing patient-facing proof is still whether a nominated intervention changes relapse activity, disability progression, imaging outcomes, symptoms, or safety tradeoffs when prospectively tested.

Drug repurposing adds another layer of understandable enthusiasm. Keramida et al. summarize the usual economic argument: repurposing may cost roughly $300 million and take 3 to 6 years, compared with about $2.6 billion and 10 to 15 years for de novo drug development [8]. Those figures explain why repurposing screens attract attention. They do not validate any one repurposed MS candidate, and the same review notes the absence of a harmonized regulatory framework for AI-generated repurposing predictions [8].

For governance purposes, the practical distinction is simple: a discovery model can be worth funding before it is worth believing as a clinical claim. It can justify experiments, not prescribing.

Response prediction is closer to the patient, so the appraisal bar changes

The Chaves et al. natalizumab-response study deserves more attention than most molecule-ranking claims because it addresses a concrete decision problem. Natalizumab is an established MS therapy. The question is not whether a model can imagine a new mechanism; it is whether patient-derived cellular imaging contains a signal that predicts who will respond. The reported 92% discovery accuracy and 88% validation accuracy are therefore clinically legible [1].

The study used machine learning on high-content imaging of CD8+ T cells. That matters because the input is not merely a billing code or a broad diagnostic label. It is closer to the biology of an individual patient. If such a model were prospectively validated and operationalized, the patient-facing consequence could be real: a clinician might avoid exposing a likely nonresponder to a high-cost, high-consequence disease-modifying therapy, or might monitor that patient differently after initiation.

That is the optimistic reading. The evidence still has to be held at its actual level. Reported validation accuracy is not the same as prospective clinical utility. It does not automatically establish calibration across health systems, robustness to sample handling, fairness across patient subgroups, or superiority over a decision pathway using standard clinical, imaging, and laboratory variables. A model that predicts response also needs a threshold policy: what happens to a patient just below the threshold, who reviews discordant cases, and whether alternative therapies produce better net outcomes.

Still, this is the kind of AI evidence that procurement and AI governance committees should separate from discovery theater. It names a drug, a patient-level sample, a measurable outcome, and a validation signal. It is not yet a deployment package, but it is closer to something a neurology service could eventually use than a ranked compound list.

Trial enrichment can change what a study sees without discovering a molecule

Falet et al. is the other study that should catch a reviewer’s eye. It does not claim to invent an MS drug. It asks whether predictive enrichment can identify patients more likely to show a treatment effect in progressive MS trials. The study analyzed six randomized controlled trials with 3,830 participants [2].

The held-out primary progressive MS anti-CD20 test set is the key result. In the top 50% predicted responders, the treatment effect was stronger and statistically significant, with a hazard ratio of 0.492, 95% CI 0.266 to 0.912, and p=0.0218. In the whole group, the corresponding result was weaker and not statistically significant, with a hazard ratio of 0.743, 95% CI 0.482 to 1.15, and p=0.179 [2]. The paper also reported a laquinimod replication with 318 participants [2].

For a trialist, that is not a cosmetic improvement. Enrichment can affect sample size planning, endpoint interpretation, inclusion criteria, and whether a biologically active therapy is missed because the trial population is too heterogeneous. In progressive MS, where treatment effects can be difficult to detect and disease trajectories vary, this is a credible place for AI to add value.

The limitation is equally important. A retrospectively discovered enrichment rule can make a past trial look more coherent. It does not by itself prove that the same rule will improve the next trial, and it does not mean the anti-CD20 therapy was AI-discovered. The model is working on patient selection and interpretation, not molecular invention.

One conflict signal also belongs in the appraisal file: a part-time DeepMind employee was among the authors [2]. That does not invalidate the work. It does mean the study should be read with the same disclosure-aware discipline that would be applied to any high-stakes predictive-enrichment method.

Disease reclassification is large and impressive, but adjacent to discovery

Ganjgahi et al. is easy to overfile under “AI drug discovery” because it uses large-scale MS trial data and produces a clinically interesting disease model. The 2025 Nature Medicine study trained an 8-state, 4-dimension reclassification on about 8,000 patients, 118,000 visits, and more than 35,000 MRIs, with validation in more than 4,000 external patients [9]. Those are unusually substantial figures for this literature.

The result matters because MS labels such as relapsing-remitting, secondary progressive, and primary progressive are clinically useful but coarse. A data-driven reclassification can reveal trajectories, inflammatory activity, disability dynamics, or imaging patterns that conventional categories blur. If incorporated into trial design, it could help sponsors enroll more homogeneous populations or analyze treatment effects with less noise.

That is still different from discovering a therapy. The model reclassifies existing trial data; it does not nominate a compound and carry it through prospective validation. Its most defensible relevance to drug development is enrichment, endpoint interpretation, and disease-state definition. That is enough to be important without being mislabeled.

The disclosure context should remain visible. The study included Novartis and Roche employees, and Novartis funded the Big Data Institute work described in the research record [9]. In a field where trial design and responder definition can affect commercial conclusions, those ties do not make the analysis unusable, but they do strengthen the case for independent external testing.

A PROBAST-style scorecard for the reviewed evidence

A prediction-model appraisal should not ask whether the AI is novel. It should ask whether the data, outcome definition, validation, and intended use support the decision being sold. Using a PROBAST/TRIPOD+AI-style lens, the reviewed MS drug-development evidence separates into three practical confidence bands.

Appraisal domainTarget and repurposing screensResponse prediction and trial enrichmentDisease reclassification
Clinical proximityLow. Outputs are targets, proteins, mechanisms, or ranked compounds.Moderate to high. Outputs can affect treatment selection or trial inclusion.Moderate. Outputs can reshape trial strata and disease-state interpretation.
Validation typeMostly in-silico, genetic, computational, or literature-supported; wet-lab confirmation is generally not shown in the reviewed record.Retrospective validation signals are stronger; Chaves reports discovery and validation accuracy, and Falet reports held-out and replication analyses.Large-scale external validation is a strength, based on the abstracted figures.
Prospective clinical utilityNot shown for an AI-nominated MS drug in the sources reviewed.Not yet established as routine-care utility; most promising for future prospective studies.Not a treatment-selection utility claim by itself.
Risk of overclaimingHigh, especially when candidate ranking is marketed as drug discovery.Moderate. The evidence is closer to decisions, but still easy to overstate as clinical deployment.Moderate. The work is powerful, but can be mislabeled as discovery.
Conflict and sponsor sensitivityDepends on study; portfolio incentives matter when a target or candidate is promoted.Falet includes a DeepMind-affiliated author; Chaves funding disclosures should be read as part of the study file.Novartis/Roche author and funding disclosures should stay visible.
Best current usePrioritize experiments and guide early portfolio review.Design prospective validation, enrichment strategies, and decision-impact studies.Improve trial stratification, disease modeling, and external validation of MS trajectories.

The scorecard changes the conversation with vendors. A company with a target-ranking engine should be asked for biological validation, not clinical adoption materials. A company with a response-prediction assay should be asked for validation cohorts, calibration, workflow burden, subgroup performance, and prospective decision-impact evidence. A group selling disease reclassification should be asked how the classes change trial design or endpoint interpretation, not whether the model has discovered a drug.

Partnership announcements are not evidence yet

Industry collaborations belong in the monitoring file, not in the evidence column. Merck and Mayo Clinic announced a research and development collaboration on February 18, 2026 to support AI-enabled drug discovery and precision medicine, with MS among the focus areas [10]. As of Q3 2026, the reviewed record did not include peer-reviewed MS drug-discovery results from that collaboration.

That is not a criticism of the partnership. Early collaborations often precede publishable results by years. The point is narrower: a named institution, a platform, and a disease area do not establish that an AI-nominated MS therapy has been validated. Procurement teams should treat these announcements as capability claims until peer-reviewed outputs, protocols, or trial records appear.

Regulatory status: do not ask the wrong clearance question

Most discovery-stage AI models in this review are not patient-facing software devices. A model that ranks targets or proposes repurposing candidates is usually part of drug-development governance, not a cleared clinical software product. Asking whether such a model is “FDA-cleared” can miss the real issue: whether the sponsor can document data provenance, model performance, bias controls, reproducibility, human oversight, and a validation path appropriate to drug development.

The question changes if the model is used to guide care for an individual patient. A natalizumab-response assay or decision-support model would need a different level of evidence, because its error can redirect therapy. A trial-enrichment model sits between those worlds: it may not prescribe to a patient in routine care, but it can decide who enters a study, how an effect is interpreted, and whether a development program continues.

That is why the regulatory and governance frame should follow intended use rather than the word “AI.” Discovery support, trial enrichment, and patient-level treatment prediction are not the same product.

What can be said without overclaiming

As of Q3 2026, the defensible evidence statement is tiered. AI-supported target discovery and repurposing screens in MS are useful for hypothesis generation and experiment prioritization. Patient-level response prediction and trial enrichment have stronger, more clinically proximal peer-reviewed support, especially in the natalizumab-response and progressive-MS enrichment studies. Large-scale disease reclassification is important for trial interpretation and disease modeling, but it remains adjacent to drug discovery.

In the sources reviewed here, there was no report of prospective clinical validation of an AI-nominated MS drug. That boundary is the difference between an honest evidence appraisal and a slide-deck claim.

References

  1. High-content imaging and machine learning of CD8+ T cells predict natalizumab response in multiple sclerosis. Nature Communications, 2025.
  2. AI-guided clinical trial enrichment in progressive multiple sclerosis. Nature Communications, 2022.
  3. Artificial Intelligence in Multiple Sclerosis: A Systematic Review. Cureus, 2023.
  4. Prioritizing drug targets and drug repurposing candidates for multiple sclerosis through integrative Mendelian randomization and protein quantitative trait loci analyses. Journal of Neuroinflammation, 2026.
  5. QSPR modeling of multiple sclerosis drug candidates using TOPSIS and WASPAS. BMC Chemistry, 2025.
  6. Systematic drug target Mendelian randomization analysis of the human plasma proteome identifies therapeutic targets for multiple sclerosis. 2023.
  7. Multi-omics analysis identifies therapeutic targets for multiple sclerosis. Proceedings of the National Academy of Sciences, 2025.
  8. Artificial Intelligence in Drug Repurposing: A Regulatory and Economic Perspective. Medicines, 2025.
  9. Data-driven reclassification of multiple sclerosis using clinical trial data. Nature Medicine, 2025.
  10. Merck and Mayo Clinic announce new research and development collaboration to support AI-enabled drug discovery and precision medicine. Merck, February 18, 2026.

Risk-of-bias scorecard

Study design
Mixed computational, retrospective, and trial-reanalysis evidence
External / prospective validation
Partial; present in response-prediction and disease-reclassification studies
Key performance metric
92% discovery / 88% validation accuracy (natalizumab response)
Overall rating
Moderate-to-high

Informational only — read the full disclaimer. This content supports procurement and research judgment, not clinical care decisions.

Submit a correction or sourcing issue

Blogarama - Blog Directory