Skip to main content
ClinicalMind logoClinicalMind

The Evidence Gap for AI in Personalized Medicine by 2030

This appraisal examines whether the RCT evidence base supports the widely cited vision that AI will transform personalized medicine by 2030. It finds a significant gap: the evidence is concentrated in narrow, non-precision applications, with zero trials by end of 2023 for multi-omics-guided therapy, pharmacogenomic dosing, or polygenic risk score screening.

Tool
AI in personalized medicine
Updated

Reviewer

Editorial Team

Editorial Team

FDA clearance status

Not discussed

A regulatory fact, reported separately from the evidence verdict.

Risk-of-bias verdict

high

The question behind AI in personalized medicine by 2030 sounds as if the literature should now be converging on a practical decision: which individualized AI systems are ready to influence care by the end of the decade? The harder answer is that the randomized evidence through the end of 2023 supports a much narrower claim. It supports some AI-assisted clinical tasks. It does not yet support the broad 2030 vision of routine AI-guided individualized medicine.

That mismatch matters because the 2030 vision is not modest. Denny and Collins projected a precision-medicine future built around seven transformations, including routine clinical genomics, AI-enabled electronic health records, big-data-driven disease taxonomies, continuous wearable monitoring, precision nutrition, diverse mega-cohorts with more than 100 million participants, and privacy-preserving data sharing.[1] A 2025 Mayo Clinic Center for Individualized Medicine perspective put the ambition even more directly, stating that genomes will be “ubiquitous in practice,” reimbursed, and accompanied by clinical decision support by 2030.[2]

Then comes the audit trail. By the end of 2023, a scoping review in The Lancet Digital Health identified 86 randomized controlled trials evaluating AI interventions in clinical practice.[3] Earlier, Lam and colleagues had screened 11,839 articles through mid-2021 and found only 39 RCTs, 0.33% of the screened literature.[4] The problem is not that no one has run RCTs of clinical AI. The problem is where those RCTs are, what they measure, and what they leave untested.

Futuristic 2030 personalized medicine vision separated from a smaller clinical research evidence base

What the RCT Map Actually Shows

The 86-trial map is encouraging if the question is whether AI can improve selected clinical processes under trial conditions. It is much less reassuring if the question is whether AI has been prospectively tested as the engine of personalized medicine.

Evidence featureWhat was found through end of 2023Why it matters for a 2030 personalized-medicine claim
Number of RCTs86 AI RCTs in clinical practiceEnough to reject the claim that AI is entirely untested, but too few and too concentrated to generalize across precision medicine
Primary endpoint direction81% reported favorable primary endpointsFavorable trial results exist, but endpoint type determines how far the conclusion can travel
Study setting63% single-centre and 92% single-countryImplementation and population transportability remain weak points
Population reportingOnly 22 of 86 trials reported race or ethnicityA field promising individualized care cannot treat missing population characterization as a minor reporting issue
Endpoint typeThree-quarters of positive results came from diagnostic-yield endpointsDiagnostic yield is not the same as better survival, morbidity, treatment selection, or patient experience
Long-term outcomesNo trial was powered for long-term mortality or disease-specific survivalThe strongest clinical claims remain outside the tested endpoint structure

Those numbers come from Han and colleagues’ review of RCTs through November 2023.[3] They are not a dismissal of the trials. They are a boundary around what the trials can prove. A positive trial of computer-aided polyp detection does not become evidence that multi-omics AI can select therapy for an individual patient. A favorable diagnostic-yield endpoint does not become evidence of reduced disease-specific mortality. A single-country study does not become proof that a model is ready for health systems with different populations, staffing patterns, incentives, and baseline risk.

Lam and colleagues’ earlier review had already pointed in the same direction. Among the 39 RCTs found after screening 11,839 articles, 77% showed AI outperforming usual care, but 90% had sample sizes below 1,000 and 49% were judged to have high risk of bias.[4] The literature was moving, but it was not yet the kind of prospective evidence base that can carry a broad clinical transformation timeline.

The Strongest Evidence Is Real, and Narrow

The most mature part of the AI RCT literature is gastroenterology imaging. In Han and colleagues’ review, 43% of all AI RCTs were in gastroenterology, nearly all testing video-based deep learning for polyp detection.[3] This is where the field has done something concrete: put an AI system into a clinical procedure, randomize care, and measure whether the clinician detects more lesions.

That deserves credit. It is also precisely where interpretation has to stay disciplined. Computer-aided detection during colonoscopy may increase detection, particularly of diminutive polyps, but the RCT evidence summarized through 2023 had not shown reductions in advanced adenoma outcomes or colorectal-cancer mortality.[3] The clinical team still has to manage the consequences: more findings, more pathology, more follow-up decisions, and the familiar question of whether a detected abnormality changes outcomes that patients would recognize as meaningful.

This distinction is often lost in strategy language. AI-assisted endoscopy is not “AI in personalized medicine” in the same evidentiary sense as AI-guided treatment selection from genomic, imaging, laboratory, and longitudinal EHR data. It is a narrower and better-tested intervention. The same caution applies to other plausible clinical AI successes, such as AI-ECG screening for low ejection fraction or insulin-dosing systems. They can be useful without proving that individualized, multi-modal, AI-directed care is ready to become routine by 2030.

Small illuminated area of AI RCT evidence contrasted with a larger untested precision medicine domain

Where the 2030 Claim Has to Stand or Fall

The most important part of the evidence map is the empty space. Among the 86 RCTs identified through the end of 2023, there were zero trials testing multi-modal AI that integrates genomics with imaging or EHR data for individualized treatment decisions; zero trials testing AI-guided pharmacogenomic prescribing against standard care; and zero trials testing polygenic-risk-score-informed screening against conventional risk assessment.[3]

These are not peripheral omissions. They are the applications closest to the promised 2030 version of personalized medicine. If the forecast is that genomes are ubiquitous in practice and paired with decision support, then pharmacogenomic prescribing is not a decorative example. It is a core workflow. A patient has a medication decision. A system interprets genotype or genomic profile with clinical context. A clinician accepts, rejects, or modifies the recommendation. The trial question is not whether the model can produce a plausible recommendation. It is whether the AI-guided loop improves care compared with standard prescribing, in the patients who would actually receive it.

The same applies to polygenic risk scores. A risk score can stratify a population in a retrospective dataset and still leave open the prospective question: does acting on that AI-supported risk estimate improve screening decisions, downstream outcomes, equity, or resource use compared with conventional risk assessment? Without randomized testing of the decision pathway, the evidence remains upstream of clinical benefit.

Multi-omics treatment selection raises the bar further. The clinical claim usually combines several hard problems at once: heterogeneous molecular data, imaging, longitudinal EHR features, missingness, treatment heterogeneity, changing standards of care, and population differences. A model that performs well on curated historical data may still fail when the decision is embedded in a clinic visit, where a specialist has to decide whether to change therapy, a pharmacist has to check safety, a payer may need justification, and the patient bears the risk of being steered toward or away from an option.

This is the point at which “AI plus genomics plus EHR” stops being a vision statement and becomes an implementation burden. Someone has to define the eligible population. Someone has to decide what happens when data are missing. Someone has to monitor override rates, adverse events, subgroup performance, and model drift. Someone has to explain why a system with an impressive validation plot should alter care before it has been prospectively compared with existing practice.

Diagnostic Yield Is Not the Same Endpoint as Individualized Benefit

A large share of positive AI RCTs measured diagnostic yield.[3] That can be an appropriate endpoint for some tools. It is not automatically a patient-centered endpoint, and it is not automatically a precision-medicine endpoint.

For a detection system, yield asks whether more targets were found. For personalized medicine, the more relevant question is usually whether the decision changed in a way that improved outcomes for the right patient. Did the system select a safer drug? Did it avoid an ineffective therapy? Did it identify a subgroup that benefited from intensified screening without increasing avoidable harm? Did it improve outcomes across the populations expected to use it, rather than only in the population from which its training and testing data were drawn?

The endpoint shift matters for governance. A procurement committee can evaluate an AI imaging tool with a trial showing higher detection during a defined procedure. A committee evaluating a pharmacogenomic AI system needs different evidence: genotype handling, medication decision logic, clinician adherence, adverse drug events, clinical outcomes, subgroup performance, and comparison with current pharmacogenomic practice or usual care. The label “AI” does not make those evidentiary requirements interchangeable.

This is also why broad AI adoption narratives are a weak substitute for trial evidence. Even if a health system is investing in data platforms, model operations, and decision support, that infrastructure does not prove that an individualized AI recommendation improves outcomes. It may be necessary groundwork. It is not the clinical result.

The Reporting Problem Is Part of the Evidence Gap

The trial count is only one problem. The inspectability of the claims is another. Han and colleagues found that only 22 of 86 RCTs reported race or ethnicity.[3] In a field that invokes personalization, that omission is not clerical. It limits the reader’s ability to judge whether the evidence applies to the populations who will be exposed to the tool.

The same concern extends to missing-data handling, external validation, subgroup performance, fairness assessment, code or model availability, and whether the trial tested the full clinical decision loop. These are the details that determine whether a model can be safely evaluated outside the site that built or first tested it. They are also where many AI studies become difficult to interpret.

TRIPOD+AI was developed to make reporting of AI and machine-learning prediction models more complete and inspectable, including a 27-item checklist for models using regression or machine-learning methods.[5] It belongs in the precision-medicine conversation because future claims will depend heavily on prediction: who is at risk, who should be screened, who should receive which drug, and who is likely to benefit or be harmed.

CONSORT-AI plays a related role for AI clinical trials, and PROBAST-style risk-of-bias questions remain unavoidable when judging prediction models. The practical point is simple: a credible precision-medicine AI trial should make clear what data were used, who was excluded, how missing data were handled, whether the model was externally validated, whether subgroup and fairness analyses were performed, and whether the endpoint matters clinically. A model cannot be governed responsibly if its basic clinical and technical assumptions are hidden.

For readers who need a broader frame for these evidence-quality issues, the adjacent appraisal on the evidence gap in machine learning for healthcare covers external validation and reporting problems across the clinical AI literature. The same logic applies here, but the stakes are sharper because precision-medicine claims often imply individualized treatment decisions rather than a narrower workflow assist.

What Can Be Said in 2026

The cleanest statement is also the most constrained: through the end of 2023, RCT evidence supports selected narrow uses of AI in clinical practice, especially diagnostic-yield applications in gastroenterology imaging, but it does not support a broad 2030 timeline for routine individualized AI-driven care.[3]

That judgment should not be inflated into a prediction that the gap can never close. The Han review stops at trials through November 2023.[3] By Q3 2026, new trials may have appeared, and fast-moving areas can mature unevenly. A serious updated evidence map would need to search the 2024–2026 literature before making a current trial-count claim.

But the burden of proof has not disappeared. Any credible 2030 forecast for AI in personalized medicine now has to account explicitly for three facts: the RCT base through 2023 was small and homogeneous; the positive results were heavily concentrated in diagnostic-yield endpoints and gastroenterology imaging; and the central precision-medicine applications—multi-modal genomics-integrated treatment selection, AI-guided pharmacogenomic prescribing, and polygenic-risk-score-guided screening—had no RCT evidence in that map.[3]

A hospital can reasonably watch these areas, build data governance, participate in prospective studies, and prepare evaluation pathways. It should be much more careful about treating the 2030 vision as a procurement assumption. Strategy papers can describe the destination. Randomized trials still have to test the road.

References

  1. Precision medicine in 2030—seven ways to transform healthcare, Cell, 2021.
  2. Imagining the Future of Individualized Medicine in 2030: The Mayo Clinic Center for Individualized Medicine Perspective, Mayo Clinic Proceedings, 2025.
  3. Randomised controlled trials evaluating artificial intelligence in clinical practice: a scoping review, The Lancet Digital Health, 2024.
  4. Randomized Controlled Trials of Artificial Intelligence in Clinical Practice: Systematic Review, Journal of Medical Internet Research, 2022.
  5. TRIPOD+AI statement: updated guidance for reporting clinical prediction models that use regression or machine learning methods, BMJ, 2024.

Risk-of-bias scorecard

Study design
randomized controlled trial
External / prospective validation
Limited
Key performance metric
Not reported in the cited evidence
Overall rating
high

Informational only — read the full disclaimer. This content supports procurement and research judgment, not clinical care decisions.

Submit a correction or sourcing issue

Blogarama - Blog Directory