Skip to main content
ClinicalMind logoClinicalMind

AI literacy mandates in healthcare lack evidence of effectiveness

Multiple authoritative bodies now mandate AI literacy for healthcare professionals, but peer-reviewed studies have yet to show that any specific training approach changes clinician behavior or improves patient outcomes. This appraisal examines the evidence base behind these policy mandates.

Tool
AI literacy training
Updated

Reviewer

Editorial Team

Editorial team, medical informatics

FDA clearance status

None

A regulatory fact, reported separately from the evidence verdict.

Risk-of-bias verdict

High

AI literacy requirements in education policy have moved faster than the evidence base behind AI literacy training. In healthcare, that distinction now matters. The AMA has called for physician education around augmented intelligence; the AAMC has issued principles for responsible AI use in medical education; the EU AI Act includes an AI literacy obligation in Article 4; and White House Executive Order 14277, while aimed at K-12 AI education rather than clinical practice, adds to the broader policy signal that AI literacy is becoming an expected institutional responsibility.[1][2][3][4]

Those signals are not trivial. Health systems are already purchasing, piloting, or governing AI tools that affect documentation, triage, imaging, risk prediction, operations, and patient communication. A workforce-readiness plan is a reasonable precaution. The problem begins when “reasonable to require” quietly becomes “this curriculum has been shown to work.” In the published healthcare education literature, that second claim is not yet supported.

Policy documents labeled AMA, EU AI Act, and Executive Order beside an empty evidence space with a question mark and stethoscope

The mandate is real; the effectiveness claim is not

A mandate can be justified before randomized trials exist. Infection prevention, cybersecurity, privacy, and emergency preparedness all include training requirements that are partly precautionary. AI literacy can fit that category: the risk is plausible, the tools are spreading, and clinicians need enough shared vocabulary to recognize when a model output should be questioned, escalated, or ignored.

But governance language often blurs three different propositions: clinicians need AI literacy; an institution should require some form of training; and a specific module, vendor certificate, or competency pathway is effective. The first two may be defensible as policy judgments. The third requires evidence that a program changes something beyond attendance records, self-reported satisfaction, or short-term test performance.

That is where the literature becomes thin. Tolentino and colleagues’ 2024 scoping review is the most direct test of the current evidence base. Across 30 AI educational programs described in 21 papers, none referenced a learning theory, pedagogy, or instructional framework. Of 19 programs eligible for evaluation assessment, only 6 reported any evaluation at all, and those evaluations stopped at Kirkpatrick level 1 or level 2: learner reaction or knowledge gain. No study measured level 3 behavior change or level 4 patient, clinical, or organizational outcomes.[5]

Kirkpatrick evaluation pyramid showing existing AI literacy studies stopping at reaction and learning, with behavior and outcomes not yet measured

Level 1 and level 2 findings are not useless. If clinicians find a course incomprehensible, irrelevant, or impossible to complete, that matters. If a module improves basic terminology or recognition of model limitations on a post-test, that may be a necessary early signal. But these measures cannot answer the questions that procurement and governance committees usually need answered: Does the training change how clinicians use AI outputs? Does it reduce unsafe automation bias? Does it improve escalation behavior? Does it change documentation quality, diagnostic decisions, patient communication, equity monitoring, incident reporting, or model-governance participation?

A completion dashboard can make an institution look prepared while leaving the operational risk untouched. A clinician can pass a quiz on bias and still over-trust a risk score during a rushed discharge. A nurse can understand that generative AI may hallucinate and still lack permission, time, or workflow support to challenge an AI-generated handoff summary. The missing levels are not academic decoration; they are where the consequence of training would actually appear.

What clinicians say they need is narrower and more practical

The demand side is not imaginary. In a 2025 mixed-methods study from Flanders, Chatzichristos and colleagues surveyed 134 clinicians and included 39 focus group participants. Only 13.8% of clinicians felt their training had prepared them for AI, and 85% wanted introductory courses.[6]

Those numbers should make health-system leaders pause, especially because they describe perceived preparation rather than abstract enthusiasm. Clinicians were not saying they wanted a branded certificate. They were saying the current educational environment had not equipped them for tools now entering clinical settings. Focus groups also found that nurses often described weak institutional support and funding for training, which is exactly where a mandate can become an unfunded transfer of responsibility to the bedside.[6]

The same study also complicates the usual policy script. Ethics training ranked lowest in interest despite the heavy emphasis that ethics receives in academic and policy discussions.[6] That does not mean ethics is unimportant. It suggests that many clinicians may first be looking for orientation: what the tool is, when it is being used, what it can and cannot do, who is accountable, when to override it, and how to document or escalate concern.

The Flanders study has clear limits. It used nonprobability sampling, had a modest survey sample, and reflects one Belgian region rather than the global clinical workforce.[6] Its value is therefore not that it settles clinician demand everywhere. Its value is that it exposes a plausible mismatch between the training institutions feel most comfortable prescribing and the support frontline staff say they lack.

Competency frameworks are converging, but measurement is lagging

Competency frameworks help governance committees name the terrain. They can distinguish basic awareness from model appraisal, workflow integration, patient communication, bias recognition, and local monitoring. Without that shared map, AI literacy becomes whatever a vendor slide deck says it is.

Ng and colleagues offer one of the more useful approaches by separating clinicians into consumer, translator, and developer tiers. The framework matters because it does not imply that every clinician needs to become a data scientist. It also links AI appraisal to evidence-based medicine, arguing that critical appraisal of AI tools operates “through the same principles as with drugs and other medical or surgical interventions.”[7]

That analogy is productive. A generalist does not need to derive every model architecture to ask whether a tool was externally validated, whether the study population matches the local population, whether the outcome is clinically meaningful, whether the comparator is appropriate, and whether performance changes after deployment. This is close to the everyday work of evidence appraisal, even if the technical vocabulary is different.

Russell and colleagues’ six-domain framework also treats “Evidence-Based Evaluation of AI-Based Tools” as a distinct competency domain.[8] That is the right emphasis for healthcare. The most consequential AI literacy failure is not that a clinician cannot define a neural network. It is that a clinician, supervisor, or committee cannot tell the difference between a model that performed well in a retrospective development paper and a tool that is safe, useful, monitored, and appropriate for the current clinical environment.

Yet frameworks do not solve the measurement problem by existing. Ang’s 2025 review states the critical gap plainly: no validated assessment instrument exists for any proposed AI literacy competency framework.[9] That means institutions may be able to say they taught a domain, but they cannot yet rely on a validated tool to say whether a clinician has achieved competence in that domain.

This distinction is often uncomfortable in committee meetings because frameworks feel actionable. They turn diffuse anxiety into domains, levels, and checkboxes. That has political and operational value. But a checklist of competencies is not the same as a validated assessment, and a validated assessment would still not be the same as demonstrated behavior change in clinical workflow.

The literature is not only shallow; it is unevenly distributed

Generalizability is another constraint. Patel and colleagues’ 2026 review found 138 North American studies on AI in medical education compared with only 5 from Latin America, the Caribbean, and Central/South America combined.[10] That imbalance matters because AI literacy is partly local: regulatory duties, available tools, language, infrastructure, documentation norms, data governance, and professional hierarchy all shape what clinicians need to know and what they can act on.

Even learner awareness data come from specific educational settings. Wood and colleagues reported in 2021 that 30% of US medical students and 50% of faculty had low AI awareness, and fewer than 15% of students felt proficient in core AI concepts.[11] Those findings support the claim that educational gaps exist. They do not establish that a particular training model closes those gaps in a durable or clinically meaningful way.

What a defensible health-system decision looks like now

For a CMIO or AI governance committee, the practical conclusion is not to wait for perfect evidence before training anyone. That would be its own governance failure. The better conclusion is to stop treating AI literacy products as if their effectiveness has already been established.

A defensible program should start by naming the local risk it is meant to reduce. A course for clinicians using ambient documentation should not be evaluated only by whether learners can define machine learning. It should examine whether clinicians review generated notes differently, whether errors are detected before signing, whether patient-sensitive statements are handled appropriately, and whether escalation pathways are used. A program for clinicians exposed to predictive risk scores should ask whether users understand intended use, contraindicated use, local validation limits, override expectations, and monitoring responsibilities.

  • Treat completion as an implementation measure, not an effectiveness measure.
  • Separate learner satisfaction from knowledge gain, and separate both from clinical behavior.
  • Require vendors or internal teams to state which competency framework they used and why.
  • Build local pre/post measures around actual AI-enabled workflows, not generic AI vocabulary alone.
  • Plan at least one behavior-level measure before calling a program validated.

The evaluation does not need to be elaborate at the start. It does need to be honest. If the institution only measures attendance and self-reported confidence, the conclusion should be limited to attendance and confidence. If it measures quiz performance, the conclusion should be limited to short-term learning. If it wants to claim safer or more appropriate AI use, it needs evidence from practice.

Procurement language should reflect the same discipline. A vendor can reasonably claim that its curriculum maps to a framework, covers high-priority domains, or has been rated favorably by participants if it has the data to support those claims. It should not be allowed to imply that the program improves clinician behavior or patient outcomes unless those outcomes have been measured.

Claim being madeEvidence needed before relying on it
Clinicians completed the trainingCompletion records and denominator definition
Clinicians liked the trainingReaction measures with response rate
Clinicians learned core conceptsPre/post knowledge or skills assessment
Clinicians use AI tools more appropriatelyObserved workflow behavior, audit data, or documented decision changes
The program improves patient or organizational outcomesOutcome measures tied to the trained workflow and plausible causal pathway

The current evidence supports action, not certainty

AI literacy is a reasonable governance priority. It is also, at present, an incompletely evaluated intervention. The peer-reviewed literature supports the existence of need, the usefulness of competency frameworks, and the urgency of preparing clinicians for AI-enabled care. It does not yet support choosing one curriculum, vendor, instructional design, or certificate as effective at changing clinician behavior or improving patient or organizational outcomes.

The cleanest path is to fund AI literacy as an iterative, locally evaluated governance intervention. Give clinicians time to complete it. Match content to the tools they actually encounter. Measure more than confidence. Require outcome plans before treating any approach as validated. Until level 3 or level 4 evidence exists, the honest claim is modest: healthcare organizations should train for AI literacy because the risk environment requires preparation, not because any published program has yet proven that it changes practice.

References

  1. Augmented Intelligence in Health Care, American Medical Association, https://policysearch.ama-assn.org/policyfinder/detail/artificial%20intelligence?uri=%2FAMADoc%2FHOD.xml-0-4261.xml
  2. Principles for Responsible Use of AI in Medical Education, AAMC, 2025, https://www.aamc.org/about-us/mission-areas/medical-education/principles-ai-use
  3. Regulation (EU) 2024/1689 laying down harmonised rules on artificial intelligence, European Union, https://eur-lex.europa.eu/eli/reg/2024/1689/oj
  4. Executive Order 14277 of April 23, 2025, Advancing Artificial Intelligence Education for American Youth, The White House, April 2025, https://www.whitehouse.gov/presidential-actions/2025/04/advancing-artificial-intelligence-education-for-american-youth/
  5. Artificial Intelligence Literacy in Medical Education: A Scoping Review, PubMed Central, 2024, https://pmc.ncbi.nlm.nih.gov/articles/PMC11294785/
  6. AI literacy among healthcare professionals in Flanders: a mixed-methods study, PubMed Central, 2025, https://pmc.ncbi.nlm.nih.gov/articles/PMC12685233/
  7. Artificial intelligence in medical education: a practical guide for medical educators, PubMed Central, 2023, https://pmc.ncbi.nlm.nih.gov/articles/PMC10591047/
  8. Competencies for the Use of Artificial Intelligence-Based Tools by Health Care Professionals, Academic Medicine, 2023, https://journals.lww.com/academicmedicine/2023/03000/competencies_for_the_use_of_artificial.13.aspx
  9. AI literacy competency frameworks for health professions education, Springer, 2025, https://link.springer.com/article/10.1007/s44217-025-00812-z
  10. Artificial intelligence in medical education in the Americas: a review, PubMed Central, 2026, https://pmc.ncbi.nlm.nih.gov/articles/PMC13157109/
  11. Medical Student and Faculty Perceptions of Artificial Intelligence in Medical Education, Journal of Medical Education and Curricular Development, 2021, https://journals.sagepub.com/doi/10.1177/23821205211024078

Risk-of-bias scorecard

Study design
Scoping review
External / prospective validation
No
Key performance metric
Not reported in the cited evidence
Overall rating
High

Informational only — read the full disclaimer. This content supports procurement and research judgment, not clinical care decisions.

Submit a correction or sourcing issue

Blogarama - Blog Directory