Verdict first: promising signal, not procurement-grade evidence
For AI in trifocal cataract surgery lens selection, the peer-reviewed record is much narrower than the marketing language around AI planning tools suggests. The direct evidence consists of one retrospective multicenter study by Cao et al. evaluating machine-learning-enhanced IOL power selection specifically for trifocal IOL implantation in axial myopia, with 139 training eyes and 48 external-validation eyes.[1]
That study is worth taking seriously. Its reported improvement is not a cosmetic model-performance statistic; it is the kind of refractive-accuracy endpoint that matters when a patient expects spectacle independence after paying for a presbyopia-correcting lens. But it is not enough, by itself, to justify broad adoption, standardization, or vendor-superiority claims for AI-based trifocal IOL power selection.
Regulatory and product status should be kept in a separate column from clinical superiority. ZEISS announced FDA clearance and CE mark availability for its AI IOL Calculator in 2024, and its product materials describe AI-supported IOL calculation functionality.[2][3] Clearance or marking can establish that a product may be marketed under a regulatory pathway; it does not establish that a tool improves trifocal cataract outcomes across the populations where a purchasing committee may want to deploy it.
| Evidence item | What it supports | What it does not support |
|---|---|---|
| Cao et al. 2024, trifocal-specific ML study | A meaningful refractive-accuracy signal in axial myopia, including external validation: 139 training eyes and 48 external-validation eyes.[1] | General adoption across routine trifocal IOL candidates, other axial-length groups, ocular comorbidities, or different clinical settings. |
| Al-Ani et al. 2026 systematic review of AI anterior-segment surgery models | A methodological benchmark: 71% of models were high risk of bias by PROBAST and only 11% used external validation.[4] | A direct answer to whether any particular trifocal IOL AI selector is clinically superior. |
| General cataract AI-formula studies | Evidence that AI formulas can perform well in broader cataract populations, such as Hill-RBF 3.0 in Natung et al. and Hill-RBF/Kane performance in extremely long eyes in Stopyra et al.[5][6] | Trifocal-specific performance unless trifocal eyes are numerous enough and separately analyzed. |
| Vendor and industry-facing claims | Context for what is being sold or promoted, including regulatory status and claimed workflow value.[2][3] | Independent proof that AI selection improves trifocal IOL outcomes. |

This appraisal is informational rather than clinical guidance. It does not choose a formula for an individual eye, replace surgeon judgment, or decide whether a particular patient should receive a trifocal IOL. The question is narrower and more operational: does the published evidence support buying or standardizing an AI-enhanced planning tool because it improves trifocal IOL power selection?
The Cao study is the one direct test, and its endpoint is clinically intelligible
Cao et al. is the study that matters most because it asks the right category of question. It did not merely test a modern IOL formula in a general cataract population and leave the reader to wonder whether a handful of presbyopia-correcting lenses were buried in the denominator. It evaluated a machine-learning approach for trifocal IOL power selection in cataract patients with axial myopia, using a retrospective multicenter design with 139 eyes for model development and 48 eyes for external validation.[1]
The reported K6 improvement is the reason this paper deserves more than a footnote. In cross-validation, the proportion of eyes within ±0.50 D improved from 71.94% to 79.14%; in external validation, it improved from 62.50% to 77.08%.[1] A half-diopter threshold is not the whole story in presbyopia-correcting cataract surgery, but it is a useful checkpoint. More eyes landing within that range means fewer conversations where the surgeon and counseling team must explain why a premium-lens result does not feel premium.
The external-validation result is especially important. Many AI models look competent inside their development dataset and then lose their advantage when moved to another site, another biometry distribution, or another documentation habit. Cao et al. at least attempts to cross that boundary, and the direction of effect in the external-validation cohort is clinically favorable.[1]
The fragility is just as visible. A training set of 139 eyes is small for a model that may be proposed as a planning standard. The external-validation set of 48 eyes is valuable but thin. The design is retrospective. The population is axial myopia only. The paper does not establish performance in typical axial lengths, short eyes, mixed corneal histories, eyes with comorbidities, or across the range of lens platforms and surgeon preferences encountered in routine premium cataract practice.[1]
That distinction matters because the outcome being sold to patients is not simply a low mean prediction error. A trifocal plan commits the clinical team to a refractive target, a dysphotopsia discussion, an expectation-management conversation, and a postoperative satisfaction risk. If an AI model has only been shown in a narrow axial-myopia cohort, the right conclusion is not that it failed. The right conclusion is that its demonstrated operating range remains narrow.
What a PROBAST-style evidence standard adds
A committee deciding whether to standardize software needs a different threshold from a reader deciding whether a paper is interesting. The relevant question is not whether the algorithm is clever or whether the reported gain is plausible. It is whether the study design, validation, population, and replication record are strong enough to support use outside the circumstances in which the model was developed.
Al-Ani et al.’s 2026 systematic review of AI models for anterior-segment surgery is useful here because it places this problem in a broader methodological pattern. Using PROBAST-oriented risk-of-bias assessment, the review reported that 71% of AI anterior-segment surgery models were at high risk of bias and that only 11% used external validation.[4]
Cao et al. does better than many AI-model papers by including external validation. That should be credited. But PROBAST-style scrutiny does not stop once the phrase “external validation” appears. The size of the validation cohort, the representativeness of the population, missing-data handling, predictor availability, outcome measurement, and the gap between retrospective performance and prospective workflow use still matter.
Under that lens, the current trifocal-specific evidence is promising but not mature. A small external cohort can show that the signal did not disappear immediately outside the development set. It cannot show stability across health systems, biometers, surgeons, lens models, ethnic and biometric distributions, or counseling practices. It also cannot tell a procurement committee what happens when the model is embedded into real scheduling, biometry review, surgeon override, and postoperative audit processes.
Replication is the missing governance fact. One group can generate a useful result; a service line usually needs to know whether another group can reproduce it. At present, no independent replicated peer-reviewed study directly confirming the Cao trifocal-specific result has been identified.
Why general AI formula studies cannot be substituted for trifocal evidence
The broader cataract literature is more encouraging than the trifocal-specific literature, but it answers a different question. Natung et al. reported that Hill-RBF 3.0 achieved 88.1% of eyes within ±0.50 D in a general cataract population; trifocal IOLs represented 10% of the study population.[5] That is useful evidence about formula performance in a wider cataract setting. It is not the same as a powered, separately analyzed trifocal IOL study.
This is not a pedantic subgroup complaint. Trifocal candidates are different from monofocal cataract patients in the consequences attached to a miss. Residual refractive error may be tolerated differently when the patient expected distance, intermediate, and near function from a premium implant. A formula that performs well across a mixed population may still need trifocal-specific analysis before it can support a trifocal procurement claim.
Stopyra et al. provides another example of useful but indirect evidence. In extremely long eyes, Hill-RBF 3.0 and Kane were reported as tied at 54.2% within ±0.50 D.[6] That kind of comparison matters for difficult biometry. It does not resolve whether AI-enhanced selection should be preferred for trifocal IOL candidates as a class.
The cleanest rule for this appraisal is simple: general cataract AI-formula performance can create plausibility, but not trifocal-specific proof. To carry a trifocal claim, the study needs enough trifocal eyes, a prespecified or clearly reported trifocal analysis, and validation that reflects the population where the tool will be used.
Industry claims should be treated as leads, not conclusions
The current commercial language around AI cataract planning often moves faster than the peer-reviewed trifocal evidence. A 2025 CRSToday article frames AI in personalized cataract care with a 20% to 30% error-reduction claim.[7] That may be directionally interesting, but unless the claim is tied to peer-reviewed trifocal IOL outcomes, it should not be treated as evidence that an AI tool improves trifocal power selection in the patients under discussion.
Vendor materials have a place in due diligence. They can identify regulatory status, intended use, interface design, input requirements, and integration features. They can also show what a company is willing to claim publicly. They are not a substitute for independent clinical validation, especially when the decision involves premium-lens outcomes where postoperative dissatisfaction is expensive in time, trust, and sometimes enhancement planning.
The ZEISS example illustrates the separation. FDA clearance and CE mark availability are relevant facts for whether a product can be considered in a regulated market.[2] Product descriptions of AI-supported IOL calculation are relevant to what the tool is designed to do.[3] Neither fact, standing alone, proves that the tool should become the standard for trifocal IOL selection.
The patient-facing stakes are real, but they do not repair the evidence gap
Trifocal selection is expectation-sensitive. Dai et al. is useful context because it places IOL choice within patient preference and shared decision-making rather than treating lens selection as a purely optical calculation.[8] That context should make evidence standards stricter, not looser. When the patient-facing promise is high, the tool used to support the plan needs a validation record that matches the promise.
A counseling team can explain uncertainty when a tool is being locally evaluated. It is harder to defend uncertainty after a health system has branded a tool as the preferred or superior AI pathway. The difference is governance: evaluation invites audit and restraint; standardization implies that the evidence is already strong enough for routine reliance.
What would be enough to change the adoption judgment?
The next useful evidence would not be another broad AI-in-cataract paper with a small premium-lens subgroup. It would be a prospective or well-controlled multicenter trifocal IOL study with enough eyes to examine relevant biometric strata, a clearly specified comparator, transparent handling of surgeon overrides, and postoperative refractive outcomes reported in a way that can be audited.
External validation should also be more than a token dataset. A validation set needs enough size and variety to show whether performance holds across sites, devices, ocular dimensions, and candidate-selection patterns. If the tool is expected to support premium-lens planning, the study should make clear which lens categories were represented and whether trifocal outcomes were analyzed separately.
For local evaluation, the governance questions are more practical: how often does the AI recommendation differ from the surgeon’s usual formula choice, who reviews discrepancies, how often the surgeon overrides the model, whether the override is documented, and whether refractive outcomes are audited by lens type. These are implementation questions, not proof of superiority, but they can prevent a promising model from becoming an unexamined default.
Procurement-grade conclusion
Cao et al. provides the only direct peer-reviewed signal in this appraisal that AI-enhanced selection may improve trifocal IOL power selection. The reported gain in eyes within ±0.50 D, including external validation, is clinically meaningful enough to monitor and possibly to justify a carefully audited local evaluation.[1]
It is not strong enough to support broad clinical adoption or procurement claims that an AI planning tool is superior for trifocal cataract surgery. The evidence is retrospective, small, limited to axial myopia, externally validated in only 48 eyes, and not independently replicated. Broader AI-formula studies support plausibility in cataract biometry, but their general populations and small trifocal representation do not answer the trifocal-specific adoption question.[5][6]
| Appraisal domain | Current status for AI-enhanced trifocal IOL power selection |
|---|---|
| Direct peer-reviewed evidence | One trifocal-specific retrospective multicenter study. |
| Sample breadth | Small and narrow: axial myopia only; 139 training eyes and 48 external-validation eyes. |
| Validation | External validation present, but limited in size. |
| Replication | No independent replicated trifocal-specific result identified in this appraisal. |
| Reported performance signal | Meaningful improvement in proportion within ±0.50 D in Cao et al. |
| Overall evidence strength | Insufficient for broad adoption, standardization, or vendor-superiority claims in trifocal cataract surgery. |
References
- Cao et al. 2024 article, Computers in Biology and Medicine.
- ZEISS Innovation Paves Way to More Personalized Patient Care With Cutting-edge AI Technology and Industry-leading Surgical Solutions, PRNewswire, August 2024.
- ZEISS AI IOL Calculator, ZEISS Medical Technology.
- Al-Ani et al. 2026 article, Ophthalmology Science, July 15, 2026.
- Natung et al. 2024 article, Indian Journal of Ophthalmology.
- AI-Based Formulas Show Good Accuracy in Calculating IOL Power for Extremely Long Eyes, American Academy of Ophthalmology, 2025.
- AI in Personalized Cataract Care, CRSToday, August 2025.
- Dai et al. 2024 article, Patient Preference and Adherence.