Ipsen’s 24 July 2026 update on BOLD is a useful place to start because the result is narrow, clinically hard, and easy to overread. Bylvay, or odevixibat, was not an AI-discovered drug. Ipsen has not yet released the full efficacy dataset, including hazard ratios, p-values, or event counts. What is public is enough for one disciplined conclusion: in a 254-infant, 19-country Phase III study in biliary atresia, the trial missed its primary endpoint of native liver survival after Kasai portoenterostomy.[1]
That endpoint matters. Native liver survival is not a biomarker of target engagement, a symptom score, or a short-run safety readout. It asks whether the infant remains alive without liver transplantation. A drug can have a coherent bile-acid mechanism and still fail that endpoint if the disease process has already moved beyond the part of biology the drug meaningfully controls.
The lesson for AI drug discovery evidence is not that this trial proves an AI-specific failure. It is that translation risk does not disappear because the mechanistic story is clean. AI systems can make candidate selection less wasteful, but they inherit the same evidentiary terrain humans built. If that terrain overrepresents positive mechanisms and underrepresents failed transfers, null endpoints, and subgroup collapses, the model’s confidence can become another polished version of the same mistake.

What Failed In BOLD Was A Clinical Endpoint
Odevixibat inhibits the ileal bile acid transporter, reducing enterohepatic bile-acid recirculation. That mechanism has already been tested in pediatric cholestatic diseases where pruritus and bile-acid burden are central clinical problems. In PEDFIC 1, odevixibat was evaluated in progressive familial intrahepatic cholestasis; in ASSERT, it was evaluated in Alagille syndrome.[2][3]
The inferential jump into biliary atresia was not irrational. It was also not automatically portable. PFIC, Alagille syndrome, and biliary atresia all sit in the pediatric cholestasis neighborhood, but neighborhood is not endpoint equivalence. Biliary atresia after Kasai involves bile duct obstruction, inflammation, progressive fibrosis, and architectural liver damage. Transplant risk is not governed only by bile-acid accumulation.

That distinction is not a semantic objection. It changes what counts as relevant evidence. A drug that improves pruritus in one cholestatic condition may still be poorly positioned to prevent transplantation in another condition where the liver’s physical architecture is already being destroyed. The endpoint moved from symptom control and cholestatic burden toward organ preservation.
BOLD therefore should not be treated as a case where a target was “disproved” in all pediatric liver disease. It is better read as a failed indication transfer to a harder, disease-specific endpoint. That narrower reading is less dramatic and more useful. It keeps the signal attached to the clinical question that actually failed.
Endpoint Proximity, Not Safety, Was The Hard Part
Early tolerability can de-risk a program, but it does not make the efficacy endpoint nearby. In BOLD, the decisive question was not whether infants could receive an IBAT inhibitor. It was whether inhibiting ileal bile-acid reuptake after Kasai would improve native liver survival. Those are different claims.
The same separation is often blurred in AI-discovery presentations. A platform may identify a plausible target, nominate a molecule rapidly, show preclinical activity, and produce an acceptable Phase I safety profile. Those achievements are real, but none of them proves that the selected mechanism is close enough to the disease-defining clinical endpoint.
For rare pediatric diseases, the temptation to accept mechanistic adjacency is understandable. Sample sizes are limited, natural histories vary, and waiting for mature outcome data can feel ethically and commercially unbearable. But the patient does not experience mechanistic adjacency. The patient either avoids transplantation, progresses, stabilizes, deteriorates, or enters another clinical pathway.
This is also where AI rare-disease claims need a tighter standard than “we found a biologically plausible match.” Drug repurposing and diagnostic acceleration can both help underserved patients, and AI may have meaningful roles in each. But for treatment selection, especially in organ-threatening disease, the relevant unit is not the elegant mechanism; it is the disease-specific endpoint. That distinction is central to evaluating claims such as AI drug repurposing for rare diseases.
The Negative Evidence Problem Is Not A Footnote
AI drug discovery systems learn from the literature, databases, assay outputs, omics resources, patents, trial registries, and proprietary datasets they are given. The problem is not that these sources contain no truth. The problem is that they are unevenly populated with success stories, publishable mechanisms, and positive associations, while many failed transfers and null experiments remain difficult to access or unusable.
A Phesi analysis reported in Drug Discovery World found that only 29.3% of more than 600,000 clinical trial protocols were linked to usable patient data.[4] That figure should be handled carefully: it does not mean every missing dataset would have changed an AI model’s prediction, and it does not measure the quality of all proprietary data. It does, however, quantify a very practical constraint. Much of the clinical-trial record is not available in a form that can train, test, or disconfirm a model.

Digital Chemistry’s white paper makes the same point from a different angle, arguing that negative data scarcity is a primary failure mode for AI-driven drug discovery.[5] As an industry white paper, it should not be treated as independent clinical proof. Its value is conceptual and operational: a model that rarely sees why plausible hypotheses failed will be poorly calibrated when another plausible hypothesis is offered.
The BOLD case makes that calibration problem concrete. The missing training signal is not simply “IBAT inhibition failed in biliary atresia.” The more useful signal is structured: which infants were enrolled, how soon after Kasai they were treated, how native liver survival behaved, whether subgroups differed, and whether bile-acid modulation was too far downstream or too small relative to fibrotic progression. Ipsen and investigators have said further analyses may provide insight into disease biology and subgroup outcomes, and the 254-patient dataset is described as the largest biliary atresia dataset assembled to date.[6]
That dataset’s value is not reputational salvage. It is disconfirming evidence, and disconfirming evidence is exactly what both human teams and AI platforms usually need after a plausible program fails. If those data remain only partially public, the next model sees the polished outline of a negative result without enough clinical texture to learn from it.
AI Platforms Can Repeat The Same Transfer Error Faster
The strongest argument for AI in early discovery is speed across a large hypothesis space. That is also the reason governance committees should ask what evidence is being accelerated. Faster nomination of disease-adjacent mechanisms is useful when followed by disciplined validation. It is dangerous when the platform converts mechanistic plausibility into decision weight before the claim has met an endpoint-relevant test.
Industry reporting has described roughly 75 AI-discovered molecules entering human trials since 2015, with no de novo AI-designed drug yet reaching FDA approval; the same reporting notes that apparently high Phase I success has not translated into clearly superior Phase II performance.[7] A separate industry analysis frames the sector as a $60 billion reality check, again emphasizing the gap between platform promise and market-access evidence.[8] These counts vary by definition, and they should be cross-checked before being used as a formal benchmark. “AI-discovered” can mean anything from algorithmic target nomination to heavily human-directed medicinal chemistry.
The direction of the evidence is still hard to ignore. No approved drug has yet shown that an AI-native discovery process reliably improves late-stage clinical success. Some AI-assisted assets, including candidates with substantial human intervention, continue to advance. That matters. The point is not that AI cannot produce a useful medicine; it is that the late-stage proof has not arrived.
Several high-profile Phase II disappointments show the recurring pattern without proving a universal law. BenevolentAI’s BEN-2293 in atopic dermatitis, Exscientia’s EXS-21546 in immuno-oncology, and Recursion’s REC-994 in cerebral cavernous malformation have all been cited as AI-associated clinical failures or stalls after encouraging preclinical or early clinical signals.[9] They are not the same disease, modality, endpoint, or platform. Their common lesson is more limited: the evidence that gets a molecule into humans often does not answer the question that determines Phase II success.
Evidence Binding Is The Governance Question
A January 2026 preprint and white paper on “Evidence-Binding Failure” describes an “Admissibility Gap”: AI-assisted drug discovery claims can look audit-ready while lacking verifiable binding between the generated claim and the evidence objects that would make the claim decision-grade.[10] Because it is not peer-reviewed, it should not be cited as settled academic consensus. But the framework names a problem that diligence teams already encounter.
A platform output that says a target is relevant to a disease is not enough. The reviewer needs to know which datasets supported that claim, whether negative trials were included, whether the endpoint in the training evidence resembles the proposed endpoint, whether pediatric and adult data were mixed, whether disease subtypes collapsed into one label, and whether the model’s confidence changes when disconfirming evidence is added.
In a case like biliary atresia, evidence binding would force uncomfortable questions early. Was the relevant endpoint pruritus, serum bile acids, time to transplant, native liver survival, or another outcome? Was the mechanism tested before or after irreversible architectural injury? Did the evidence come from PFIC and Alagille syndrome, or from biliary atresia itself? Were failed attempts in related cholestatic diseases represented, and were they machine-readable?
These are not anti-innovation questions. They are procurement, investment, and trial-design questions. A committee asked to approve an AI-discovered candidate does not need a more beautiful pathway diagram. It needs traceability from claim to evidence, and it needs to know where the evidence stops.
What BOLD Should Change In AI Drug Discovery Appraisal
The most useful response to BOLD is not to say that AI would have predicted the outcome. There is no public basis for that claim, and the trial was not AI-driven. A model trained mostly on positive cholestasis literature might have liked the same mechanistic story humans liked. Without disease-specific negative evidence and endpoint-grounded validation, it might simply have arrived at the same overextension with more computational authority.
For AI drug discovery claims, BOLD supports a practical appraisal sequence:
- Separate target plausibility from endpoint proximity.
- Ask whether the training and validation evidence includes failed indication transfers, not only successful mechanisms.
- Check whether disease labels hide biologically distinct subgroups, especially in pediatric rare disease.
- Require traceable evidence-object binding for claims that influence trial design, investment, or procurement.
- Downgrade platform claims that report Phase I acceleration without showing Phase II or endpoint-specific validation.
Pediatric rare disease also has a diagnostic-data problem that should not be confused with therapeutic proof. AI may help shorten diagnostic odysseys by recognizing patterns across sparse clinical records; that is a different evidence task from predicting whether a drug will preserve an infant’s native liver. The distinction matters when reading claims about AI in pediatric rare disease diagnosis.
BOLD does not prove AI drug discovery cannot work. It shows why plausibility plus early safety is not decision-grade evidence. Before an AI platform is treated as more than a hypothesis-generation engine, it should show access to disease-specific negative data, validation against clinically proximate endpoints, and traceable binding between each claim and the evidence that supports or weakens it.
References
- Ipsen provides update on Phase III BOLD trial in biliary atresia, Ipsen, 24 July 2026.
- Efficacy and safety of odevixibat in patients with progressive familial intrahepatic cholestasis (PEDFIC 1): a randomised, double-blind, placebo-controlled, phase 3 trial, PubMed.
- Efficacy and safety of odevixibat in patients with Alagille syndrome (ASSERT): a phase 3, double-blind, randomised, placebo-controlled trial, PubMed.
- AI not solving clinical trial flaws, report says, Drug Discovery World, 2026.
- The Hidden Value of Failure: Why Negative Data is Critical for AI-Driven Drug Discovery, Digital Chemistry.
- Ipsen’s Bylvay fails to improve native liver survival in biliary atresia Phase III trial, AllSci.
- First AI-designed drugs fall short in the clinic following years of hype, Endpoints News.
- AI Drug Discovery’s $60 Billion Reality Check: Hype, Failures, and the Market Access Blindspot, Loonbio.
- AI drug discovery failure record, note.com.
- Evidence-Binding Failure in AI-Assisted Drug Discovery Pipelines: The Admissibility Gap Between Plausibility and Decision-Grade Support, ResearchGate, January 2026.