Most healthcare AI accessibility reviews still answer the wrong question first. They can tell you whether a model looks accurate, whether a page passes a scanner, or whether a vendor says the interface is inclusive, but they often do not show whether a disabled patient can actually use the system in a clinic, with the assistive technology they rely on. Waters (2026) is strongest because it tries to turn that gap into a testable process instead of a vague obligation [1].

What Waters adds to AI accessibility reviews
Waters frames accessibility TEVV around four linked components: red teaming, model testing, field testing, and usability testing [1]. That matters because each layer catches a different kind of failure. Red teaming pushes on adversarial or edge-case breakdowns. Model testing asks what the system does across disability-relevant subgroups. Field testing moves the review into clinical conditions instead of a clean demo environment. Usability testing checks whether the person using the system can complete the task with the tools, time pressure, and constraints that matter in practice [1].
The framework is more useful than a generic accessibility checklist because it ties those layers to measurable targets. Waters proposes seven quantifiable metrics, including an Inclusive Accuracy Rate of at least 95%, an Accessibility Disparity Index of 0.05 or less, and an Assistive Technology Compatibility Score of at least 90% [1]. The point is not that these thresholds are magic numbers. The point is that they force evaluators to say what is being measured, for whom, and under what conditions, which is exactly where many accessibility reviews become vague.
That structure is also where the standards matter. Waters roots the framework in ISO 9241-210, WCAG 2.2, EN 301 549, ISO/IEC 25010, and the NIST AI Risk Management Framework, so accessibility review is treated as part of a broader TEVV process rather than as a decorative add-on [1]. For clinical evaluators, that is the important shift: the question is no longer only whether the interface passes a scan, but whether the system has been tested in a way that can support a real decision about deployment.

Why the framework fills a real gap
A systematic review by Chemnad and Othman found that AI accessibility research from 2018 to 2023 was disproportionately focused on visual impairments, while motor and cognitive disabilities received much less attention [2]. That imbalance matters because a system that looks accessible to one group can still fail another group in ways a scanner will never catch. It also means a passing review may simply reflect which disability categories were easiest to study, not which ones are easiest to serve.
Fuglerud and colleagues add the practical caution. Their work shows that AI tools can be decent at spotting technical WCAG violations such as color contrast problems, but they struggle with contextual accessibility judgments and can produce enough false positives that human review still has to clean up the result [3]. That is the core limitation for using AI in accessibility reviews for people with disabilities: automated checks can confirm that a box was ticked, but they do not reliably tell you whether the box matters in use.
Seen together, those findings explain why surface-level compliance can look reassuring while leaving disabled users to absorb the failure later. The system may satisfy a procurement checklist, but the patient still has to navigate reachability, input methods, screen-reader behavior, timing, or clinic workflow on the day of care. Waters is trying to make those hidden conditions part of the evaluation itself rather than a post-deployment complaint.
Where the framework is still a proposal
The most important limitation is also the most obvious one: Waters is a hypothesis-and-theory framework, not an empirical validation study [1]. The article’s illustrative applications, such as clinical assistant or patient-screening scenarios, show how the method could be used, but they do not prove that the thresholds predict real accessibility outcomes. A framework can be rigorous and still be untested.
That means the metrics should be read as disciplined prompts for review, not as substitutes for review. Inclusive Accuracy Rate, Accessibility Disparity Index, and Assistive Technology Compatibility Score are useful because they make accessibility legible to clinical researchers and regulators, but they still depend on the choice of test set, the assistive technologies included, the tasks selected, and the people invited into the evaluation. If the wrong users are excluded, a precise score can still hide a weak test.
The same caution applies to any claim that accessibility can be fully automated. Even strong tooling is still best understood as a filter for obvious technical defects, not as a complete judgment about whether disabled patients can actually use a healthcare system. Waters is valuable because it does not confuse those two things.
The broader evidence gap in healthcare AI is discussed in The Evidence Gap in FDA-Cleared AI Medical Devices, and the accessibility framework belongs in that same conversation: it narrows one of the places where clearance or procurement can look stronger than real-world use.
References
- Waters, "AI testing, evaluation, verification and validation for accessibility: a comprehensive framework," Frontiers in Digital Health, Feb. 2026, https://pmc.ncbi.nlm.nih.gov/articles/PMC12980396/
- Chemnad and Othman, systematic review on AI accessibility research coverage, 2024, https://pmc.ncbi.nlm.nih.gov/articles/PMC10905618/
- Fuglerud et al., study on AI support for accessibility testing, 2024, https://pubmed.ncbi.nlm.nih.gov/39560272/
Comments
Join the discussion with an anonymous comment.