
The evidence on AI in medical education is positive, but the defensible claim is much narrower than the usual headline. Randomized and quasi-experimental studies often report better short-term test or skills scores when AI tools are added to teaching. They do not yet show that AI broadly improves medical learning, changes clinical behavior, or produces more competent clinicians.
The best recent randomized synthesis is encouraging on its face: 20 RCTs with 1,413 learners found improved knowledge outcomes and clinical skills after generative AI was integrated into medical education. But the same review found no trial at low overall risk of bias, no sham-AI or attention-matched control, and only low GRADE certainty for knowledge outcomes.[1] That is the basic tension in the literature: the signal is real enough to study seriously, but not clean enough to support curriculum-wide claims.
| Appraisal question | Current answer |
|---|---|
| Regulatory status | No FDA clearance angle applies; these are educational interventions, not cleared clinical devices. |
| Study designs | Mostly RCTs and quasi-experimental education studies, often testing AI as a supplement rather than as the full instructional intervention.[1][2] |
| Evidence quality | Low to moderate at best in the strongest RCT synthesis; low to very low in a broader health-professions review.[1][2] |
| Generalizability | Limited by small, single-center studies and geographic concentration, especially China-based trials.[1][3] |
| Risk of bias | Material: blinding is rare or unreported, allocation methods are often weakly described, and attention-matched controls are generally absent.[1][3] |
| Outcome level | Outcomes remain at the learning or performance-in-study level; evidence for behavior change or patient/clinical-system outcomes is not established.[2] |
The pooled RCT results are favorable, then immediately become complicated
Wang et al. is the review that deserves the most weight because it restricts the question to randomized controlled trials and short-term learning outcomes. Across 20 RCTs, generative AI was associated with a knowledge standardized mean difference of 0.99, with a 95% CI from 0.49 to 1.49. Clinical skills showed a similar pooled effect, with an SMD of 1.09 and a 95% CI from 0.81 to 1.38.[1]
Those are not trivial effects. If taken at face value, they suggest that structured AI-supported practice can help learners perform better on near-term assessments. The important restraint is that the review itself does not give permission to treat those pooled estimates as settled. The knowledge outcome had low GRADE certainty, and the prediction interval ran from −0.57 to 2.56, meaning a future trial in a different setting could plausibly show little benefit, no benefit, or a large benefit.[1]
The missing comparator matters. None of the 20 RCTs used a sham-AI or attention-matched control.[1] In education trials, that is not a technical quibble. A student who receives immediate, interactive, novel feedback may do better than a student assigned to usual materials because of feedback frequency, time on task, novelty, or tutor-like attention—not necessarily because the model generated uniquely better instruction.
This distinction is especially important for curriculum committees. A study can truthfully show that an AI-supported learning session improved a post-test and still leave unanswered whether a non-AI tutor, a structured worksheet, a retrieval-practice module, or a faculty-designed feedback script would have done the same thing if given equal time and interactivity.
Other meta-analyses widen the effect range rather than settling it
Peng et al. reported much larger effects in a China-specific meta-analysis of 12 studies involving 824 undergraduate medical students. Exam scores improved with an SMD of 2.06, with a 95% CI from 1.35 to 2.76, and satisfaction favored AI-assisted education with an odds ratio of 5.80. But heterogeneity was very high for exam scores, with I² at 94%; most studies did not report allocation concealment or blinding; and sample sizes ranged from 18 to 120.[3]
That pattern is familiar in early education technology research: the effect looks dramatic, the learner experience looks favorable, and the methods leave several doors open. Satisfaction is also not competence. It is useful for implementation planning, but a satisfied learner may be responding to convenience, novelty, reduced anxiety, or a sense of support rather than to deeper learning.
Pillai et al. reported more modest randomized effects across 14 RCTs and 1,116 learners: knowledge SMD 0.36, skills SMD 0.78, and satisfaction SMD 0.97. The authors also acknowledged likely publication or reporting bias and high heterogeneity.[4] That result is probably closer to the tone committees should use: plausible benefit, stronger for skills and satisfaction than for knowledge, and still vulnerable to selective reporting and uneven study quality.

The broader health-professions review applies the brake
Feigerlova et al. is useful because it refuses to smooth weak studies into a more confident pooled answer. The review included 12 studies in health professions education, all single-center, with sample sizes ranging from 4 to 180. Only 3 were RCTs; 2 of those were judged high risk of bias and 1 had some concerns. The 7 quasi-experimental studies were all judged at serious risk of bias.[2]
The authors rated the evidence low to very low by GRADE and judged meta-analysis inappropriate.[2] That choice is methodologically important. Pooling can create the appearance of precision when the included studies differ too much in intervention, learner group, outcome, assessment timing, and bias structure. A forest plot does not rescue a body of evidence if the underlying comparisons are unstable.
The same review found that no study reached Kirkpatrick levels 3 or 4—that is, no evidence that AI education interventions changed learner behavior in practice or improved downstream organizational or patient outcomes.[2] This does not invalidate short-term gains. It places them where they belong: at the learning-outcome level.

Answering medical questions is not the same as improving medical learning
Some of the most visible AI-in-medical-education papers do not test education at all. Kung et al. showed that ChatGPT performed at or near the passing threshold on the USMLE, with open-ended accuracy of 75.0%, 61.5%, and 68.8% across the examined components and an approximate 60% passing threshold.[5] That is a model-capability finding. It says the model can answer many exam-style medical questions. It does not show that students learn more, retain more, reason better at the bedside, or transfer knowledge into patient care after using it.
BEME Guide No. 84 made the same problem harder to ignore. Gordon et al. catalogued 32 studies comparing large language models against examinations and concluded that further studies of that form appear unjustified unless they bring a new perspective.[6] The field does not need many more demonstrations that a model can sit for a test. It needs studies that assign learners to plausible educational strategies and then measure what changes.
That is also why procurement and governance teams should separate “the model knows enough medicine to be useful” from “the product improves medical education.” The first may be necessary. It is not sufficient.
The bedside RCT shows where the controlled-practice evidence may stop
Saloojee et al. is a helpful counterweight because it put AI into a real clinical encounter rather than a quiet study session. In a randomized trial of 73 final-year students across four hospitals, students used ChatGPT powered by GPT-4o during real-patient bedside examinations. Performance did not improve: scores were 66.7±9.9 with ChatGPT and 68.2±7.8 without it, with an adjusted p value of 0.21.[7]
The process findings are as important as the null score result. Thirty-seven percent of students reported distraction, and 45% said AI slowed the encounter. Prior academic performance predicted examination scores, with p=0.04; AI use did not.[7]

This trial should not be read as disproving the positive meta-analyses. It tested a different use case. The meta-analytic signal is strongest for structured, supplemental learning in controlled settings. A real bedside examination adds patient interaction, time pressure, etiquette, cognitive load, and the need to maintain clinical flow. A tool that helps during deliberate practice can still interfere when inserted into a patient-facing performance.
For clerkship directors, that boundary matters more than the average effect size. The question is not whether a student can get useful practice from an AI tutor on Tuesday evening. Many probably can. The question is whether the same tool improves observable clinical performance when the student is examining a patient, organizing findings, responding to uncertainty, and being assessed in real time. That evidence has not arrived.
Publication bias likely inflates the field-level estimate
A late caution comes from a non-peer-reviewed OSF preprint meta-meta-analysis. It examined 1,840 effect sizes across 67 meta-analyses of AI and learning and found severe publication bias. After publication-bias adjustment, the standardized mean difference was 0.196, with a 95% credible interval from 0.000 to 0.323 and a prediction interval from −1.521 to 1.908. The authors wrote that “broad claims of generalized learning gains … appear premature.”[8]
Because this is a preprint, it should not be treated as the final adjudication of AI education effects. It is still a useful stress test. If the field’s apparent effect shrinks sharply after correcting for small-study and publication-bias patterns, then committees should be wary of treating the published pooled estimates as the expected local effect in their own curriculum.
This is the same type of evidence-gradient problem seen in other early AI appraisal areas, where high technical promise does not automatically translate into proven educational or clinical value. For readers comparing maturity across domains, the broader ClinicalMind appraisal on how AI in medicine and healthcare performs in 2026 uses a similar distinction between measured task performance and implemented value.
What the studies show, claim by claim
| Claim | What the evidence supports | Judgment |
|---|---|---|
| AI improves short-term knowledge scores in controlled education studies. | Supported, with pooled RCT effects ranging from modest to large and low certainty for knowledge in the most relevant RCT review.[1][4] | Plausible benefit. |
| AI improves short-term clinical skills assessments. | Supported in pooled RCTs, including SMD 1.09 in Wang et al. and SMD 0.78 in Pillai et al.[1][4] | Promising, still method-limited. |
| Students like AI-assisted medical education. | Supported by satisfaction findings, including OR 5.80 in Peng et al. and SMD 0.97 in Pillai et al.[3][4] | Useful for implementation, not evidence of competence. |
| AI teaches medicine better than traditional instruction. | Not established because most studies do not use sham-AI or attention-matched controls, and many are small, single-center, and unblinded.[1][2] | Overclaim. |
| AI improves real clinical performance. | Not established; the bedside RCT found no performance improvement and reported distraction and slowing during real-patient exams.[7] | Unproven. |
| AI education tools improve clinical behavior or patient outcomes. | Not established; reviewed outcomes have not reached Kirkpatrick levels 3 or 4.[2] | No current evidence base. |
| Published meta-analytic effects can be used as expected local effect sizes. | Unsafe; heterogeneity, geographic concentration, small studies, and likely publication bias make direct translation unreliable.[1][3][8] | Use only as hypothesis-generating. |
The narrow conclusion is the honest one: structured, supplemental AI use can plausibly improve short-term knowledge and skills performance in controlled medical-education settings. Broad curriculum-wide improvement is unproven. Real clinical-performance benefit is not established. The evidence base remains limited by small and single-center studies, frequent lack of blinding, weak or absent attention-matched comparators, geographic concentration, low-to-very-low certainty ratings in key reviews, and publication bias.
References
- The impact of integrating generative artificial intelligence into medical education on short-term learning outcomes: a systematic review and meta-analysis of randomized controlled trials — BMC Medical Education, 2026.
- A systematic review of the impact of artificial intelligence on educational outcomes in health professions education — BMC Medical Education, 2025.
- Effectiveness of AI-assisted medical education for Chinese undergraduate medical students: a meta-analysis — BMC Medical Education, 2025.
- Meta-analysis of randomized controlled trials comparing outcomes of AI-based teaching versus traditional-based teaching in medical education — Postgraduate Medical Journal, 2026.
- Performance of ChatGPT on USMLE — PLOS Digital Health, 2023.
- A scoping review of artificial intelligence in medical education: BEME Guide No. 84 — Medical Teacher, 2024.
- AI at the bedside: Randomised controlled trial of ChatGPT's impact on student performance in real-patient clinical exams — Medical Teacher, 2026.
- Effect of Artificial Intelligence on Learning: A Meta-Meta-Analysis — OSF preprint.