This evidence appraisal is for research, governance, and procurement review. It is not clinical guidance for diagnosing or treating an individual athlete. For readers asking what studies show about AI in sports injury diagnosis and recovery, the diagnosis half needs to be separated from the recovery half immediately: this article is about imaging diagnosis, chiefly knee MRI and related musculoskeletal imaging, not injury-risk prediction or return-to-sport forecasting. Those adjacent questions are covered separately in How Reliable Is AI for Sports Injury Prediction? and AI Sports Injury Recovery Prediction Lacks Validated Clinical Evidence.
The central split is simple and often blurred in vendor conversations: diagnostic-accuracy literature can show that a model detects a defined imaging finding in a defined dataset; FDA clearance answers a regulatory question about whether a device may be marketed for an intended use. In orthopaedic AI, that difference matters. A 2026 audit of FDA-cleared orthopaedic AI/ML medical devices found 70 cleared devices as of February 2025, but only 8.6% had been validated through a formal prospective clinical trial; 68.6% used retrospective validation, and 22.8% had no reported clinical testing. The same review notes that more than 95% of cleared AI/ML medical devices go through the 510(k) pathway, which is based on substantial equivalence rather than an independent FDA certification of diagnostic performance. [1]

What is genuinely interesting about AI reading a knee MRI
A knee MRI is one of the more plausible places for diagnostic AI to look useful. The task can be made narrow: detect an ACL tear, identify a meniscus abnormality, flag a fracture, segment an anatomic structure, or measure alignment. That is different from asking a model to understand “athlete recovery” or predict the social, physiologic, surgical, and rehabilitation factors behind return to play.
The constraint matters. A discrete imaging target can be labeled, adjudicated, split into training and test sets, and compared against radiologist interpretation or surgical reference where available. A model that performs well on that kind of task is not a toy just because it was studied retrospectively. Retrospective work is often how imaging AI starts: it lets investigators test whether the visual signal is learnable before a hospital exposes clinicians to the software in a live workflow.
But the label on the evidence decides how far the result travels. A high-performing ACL model in a research dataset does not automatically become a clinically dependable ACL reader in a sports-medicine service. The model may have been tested on curated exams, on images from a limited set of scanners, with exclusions that remove the messiest cases, or against a reference standard that differs from local practice. It may also be an academic model that is not the same software version, user interface, triage behavior, or intended-use claim being sold to a hospital.
That is why unverified accuracy snippets are a bad foundation for purchasing decisions. The knee-MRI literature includes promising reports on ACL and meniscus detection, but where the full methods, denominators, exclusions, external validation status, and workflow setting cannot be checked, the numbers should be treated as leads for due diligence rather than settled evidence. The useful question is not “did someone publish a high accuracy?” It is “was the same tool, for the same intended use, tested in a population and workflow close enough to ours?”
FDA clearance is not the same question as diagnostic accuracy
FDA clearance is still meaningful. It is not a sticker a vendor prints at home, and it should not be dismissed. The problem is the slide that lets “FDA cleared” stand in for “prospectively proven to improve diagnosis in our clinic.” Those are different claims.
The FDA maintains a public list of AI-enabled medical devices, and that list should be checked directly because entries change as new devices are added or updated. [2] But the presence of a device on that list does not tell a radiology lead whether the tool reduced misses in a live MSK MRI queue, whether it changed reporting time, whether orthopaedic surgeons trusted the output, or whether false positives created new follow-up work.

This distinction is familiar from broader medical-AI appraisals: regulatory clearance and clinical validation sit close enough to be confused, but not close enough to be interchangeable. The same clearance-versus-evidence issue is discussed in FDA-Cleared ML in Healthcare: Where’s the Clinical Evidence? and The Evidence Gap in FDA-Cleared AI Medical Devices. Orthopaedics deserves its own look because the device landscape is not simply a miniature version of all medical AI; its dominant functions and validation patterns are specific.
What the cleared orthopaedic AI market actually substantiates
Lee et al.’s audit is useful because it stops the discussion from floating around generic enthusiasm. It asks what has actually been cleared in orthopaedic AI/ML and what evidence was reported for those devices. As of February 2025, the authors identified 70 FDA-cleared orthopaedic AI/ML medical devices. Spine accounted for 42.9% of the cleared-device landscape, and hip/knee accounted for 20%. [1]
| Question | What Lee et al. found |
|---|---|
| How many FDA-cleared orthopaedic AI/ML devices were identified? | 70 devices as of February 2025. [1] |
| How many had formal prospective clinical-trial validation? | 8.6%. [1] |
| How many were validated retrospectively? | 68.6%. [1] |
| How many had no reported clinical testing? | 22.8%. [1] |
| Which anatomic areas dominated? | Spine 42.9%; hip/knee 20%. [1] |
| What function led recent new-device clearances? | Surgical planning represented 70.6% of new-device functions in 2022–2024. [1] |
That mix should temper how sports-injury diagnosis is discussed. The cleared orthopaedic AI market is not mainly a catalog of autonomous knee-MRI readers for ACL tears. In the more recent clearance period studied by Lee et al., surgical planning was the leading new-device function, representing 70.6% of 2022–2024 new-device functions. Deep learning did become more common in recent clearances, accounting for 57.3% of clearances during 2022–2024, but algorithmic sophistication is not the same as prospective proof in a sports-imaging workflow. [1]
There is also improvement in the landscape, and it is worth saying so plainly. Devices with no reported clinical testing fell from 62.2% in 2017–2019 to 19.7% in 2022–2024. [1] That is not nothing. It suggests the field is moving away from the thinnest evidence submissions. A procurement committee should still ask what kind of testing replaced the missing testing, because a retrospective single-institution validation and a prospective multi-site clinical trial do not carry the same operational meaning.
Retrospective validation deserves credit, but not a promotion
It is lazy to treat retrospective validation as worthless. In imaging AI, it can answer a real technical question: can the model identify the labeled abnormality in stored exams under the study’s conditions? For a narrow finding such as an ACL tear, that is a meaningful first screen. If a model cannot perform there, it has no business being piloted in a busy radiology department.
The problem starts when retrospective validation is promoted into a live-clinic claim. Prospective use introduces different failure modes: unreadable or atypical scans, changed acquisition protocols, local scanner variation, time pressure, interface friction, alert fatigue, and disagreement about who is responsible when the AI output and the radiologist’s impression diverge. Those are not philosophical objections. They are exactly the things a radiology group, sports surgeon, or CMIO has to manage after go-live.
A formal prospective clinical trial is not the only possible source of useful evidence, but its scarcity in cleared orthopaedic AI should slow the procurement conversation. In Lee et al.’s audit, only 8.6% of cleared orthopaedic AI/ML devices had that level of validation. [1] For a tool being positioned as clinically ready, a department should know whether it has been tested prospectively at all, and if not, what evidence is being offered instead.
The vendor packet needs to match the workflow, not just the anatomy
For a sports-medicine imaging purchase, “knee” is too broad. A knee MRI tool may be intended for triage, detection, measurement, segmentation, report drafting, surgical planning, or quality control. Each use creates a different risk. A false negative in a triage tool may delay review. A false positive in a detection tool may generate unnecessary scrutiny or follow-up. A planning tool may affect preoperative measurement rather than the diagnostic call itself.
- What exact finding or task is the tool cleared and marketed for: ACL tear, meniscus abnormality, fracture identification, alignment measurement, segmentation, surgical planning, or another function?
- Was the evidence retrospective, externally validated, multi-site, prospective, or post-market? The answer should be tied to the marketed product, not to a related academic model.
- Was the validation population similar to the local population: adolescents, collegiate athletes, elite professionals, older recreational athletes, postoperative knees, or mixed orthopaedic patients?
- What scanners, MRI protocols, image planes, and sequence types were included or excluded?
- What reference standard was used: radiologist consensus, surgical findings, chart review, or another label source?
- How does the output enter the radiologist’s worklist or report? A silent background measurement tool and an interruptive abnormality alert need different governance.
- Who reviews discordance between the AI result and the human interpretation, and where is that review documented?
These questions are not designed to make adoption impossible. They are designed to keep evidence attached to intended use. A hospital can reasonably pilot a tool with retrospective support if the risk is low, the workflow is supervised, the local validation plan is explicit, and the claims remain modest. The same evidence would be inadequate if the tool is being sold as a replacement for specialist interpretation or as a decisive diagnostic gatekeeper.
What about safety signals and recalls?
Lee et al. reported no recalls among orthopaedic AI/ML medical devices, compared with roughly 10% across all FDA-cleared AI/ML medical devices. [1] That is reassuring only within limits. The authors also note the short post-clearance follow-up window, which makes it premature to treat zero recalls as proof of durable safety. Absence of recalls is not the same as evidence that a device improves diagnosis, reduces missed injuries, or fits the daily cadence of an MSK radiology practice.
For governance committees, the recall finding belongs in the safety file, not the efficacy file. It may reduce one kind of concern, but it does not answer whether the device performs well on local knee MRIs, whether radiologists over-rely on it, or whether it creates downstream work for orthopaedics and sports medicine.
Where this leaves AI for sports-injury imaging diagnosis
The most defensible reading is neither dismissal nor procurement-by-demo. AI can be technically impressive on narrow musculoskeletal imaging tasks, and knee MRI is exactly the kind of bounded domain where deep learning may produce useful diagnostic assistance. That is the part worth being curious about.
The cleared orthopaedic device landscape, however, supports a more cautious operational conclusion. The market includes many devices that were not prospectively validated, and the dominant recent functions are not simply sports-injury diagnosis. FDA clearance should be treated as a regulatory starting point, not as a substitute for appraising clinical evidence tied to the actual product, intended use, dataset, and workflow.
For a radiology or sports-medicine department, the practical answer is: ask for the evidence label before admiring the number. Retrospective, external, prospective, post-market, same tool, same version, same task, same workflow—those words decide whether a promising model result is ready for supervised local piloting or still belongs in the research file. This article is intended to support that evidence review and procurement judgment, not to guide individual diagnosis or treatment.
References
- FDA-Cleared Artificial Intelligence Medical Devices in Orthopaedic Surgery, JAAOS Global, 2026.
- Artificial Intelligence-Enabled Medical Devices, U.S. Food and Drug Administration.