Procurement verdict: credible triage, not autonomous diagnosis
For teams evaluating AI-assisted glaucoma screening in optometrist-referral pathways, the answer is no longer that the field has only retrospective validation. Prospective evidence now exists, and it is useful. It shows that AI-assisted glaucoma screening can make referral decisions more consistent and reduce unnecessary specialist referrals. It also shows that performance can drop sharply when the model leaves curated images and enters routine primary-care conditions.
The practical verdict is narrower than many procurement decks imply: AI glaucoma screening is supportable as a standardized triage layer, especially where ophthalmology services are receiving low-yield referrals, but not as an autonomous diagnostic substitute. In the Portugal Glaucoma PORT trial, MONA-GLC referred 10% of screened participants compared with 18% referred by an expert panel, with 78% sensitivity and 95% specificity.[1] In the Australian GP pragmatic trial, real-world sensitivity was 65%, down from 95.6% in lab settings, and only 66.9% of participants had analyzable images.[2]

That gap matters more than the usual “AI works” or “AI fails” framing. Referral pathways fail in ordinary places: the retinal photograph is hazy, the older patient has cataracts, the optometrist assumes someone else will chase follow-up, the specialist clinic receives another false-positive referral, or the highest-risk patient never enters the AI denominator because the image is ungradable. A safe pathway has to account for those failures before it counts referral reduction as success.
What the Portugal PORT trial supports
The Portugal trial is the cleaner procurement signal. It tested MONA-GLC in primary care with 671 screened participants and compared AI referral behavior with an expert panel. The headline result is operationally attractive: the AI referred 10% of participants, while the expert panel referred 18%. The model’s reported sensitivity was 78%, specificity 95%, and positive predictive value 53%.[1]
That combination is exactly why ophthalmology services will pay attention. A lower referral rate with high specificity means fewer people sent into specialist queues without clear need. The 53% positive predictive value is also meaningful in a screening context, because many existing glaucoma referral pathways generate substantial false-positive burden before a specialist confirms disease.
The economic analysis is promising but should not be lifted out of its setting. The PORT investigators reported an incremental cost-effectiveness ratio of €1,725 per QALY at 1% glaucoma prevalence, and cost-saving dominance at prevalence of 2% or higher.[1] Those figures are useful for understanding the direction of value, but they are not a US budget impact model. Software licensing, imaging workflow costs, optometry reimbursement, specialist visit pricing, and medicolegal expectations can move the result substantially. Because full methodological detail may not be equally accessible to every buyer, teams should verify local assumptions rather than relying only on headline cost-effectiveness numbers.
The PORT result therefore supports procurement under a specific claim: AI may reduce unnecessary referrals while preserving a clinically useful detection rate in a structured primary-care screening pathway. It does not prove that the same model can be dropped into every optometrist referral stream with the same sensitivity, the same image quality, or the same downstream attendance.
The Australian pragmatic trial is the stress test
The Australian GP pragmatic trial is less flattering and more important for implementation. It tested automated retinal photography and AI glaucoma screening in routine primary care, where image quality, patient comorbidity, operator variation, and follow-up behavior are not controlled away. In that environment, sensitivity was 65%, compared with 95.6% in lab settings.[2]
The most consequential number is not only the sensitivity. It is the denominator. Only 66.9% of participants had images that were analyzable by the AI.[2] That means roughly one-third of the screened population did not receive the AI assessment that a procurement dashboard might otherwise treat as the intervention.

The excluded group was not random noise. Participants with ungradable images were older, with a mean age of 71.4 years compared with 64.8 years for those with gradable images, and were more likely to have cataracts, 40.9% versus 19.1%.[2] In glaucoma screening, that is not a minor technical footnote. It is a pathway safety issue because age and ocular comorbidity are exactly where risk tends to concentrate.
The trial authors also discussed possible generalizability issues because the model had been trained on a Chinese population, while the deployment population was Australian.[2] That does not prove poor performance was caused by ethnicity or training-set mismatch; image quality and patient comorbidities were also plausible contributors. It does mean buyers should ask for local validation rather than accepting aggregate performance from a different population as sufficient.
The follow-up results are another warning against treating AI screening as a self-contained diagnostic event. Only 37.5% of AI-referred patients attended specialist follow-up, and only 16.7% of those who attended had confirmed glaucoma.[2] A model can generate a referral recommendation, but the pathway still has to get the patient to the specialist, close the loop, and learn from false positives and false negatives.
The same trial still found signal where the current system misses disease: AI detected referable glaucoma in 11.2% of previously undiagnosed patients who had seen an eyecare provider in the prior 12 months.[2] That is why the Australian evidence should not be read as a rejection of AI. It is a rejection of assuming that good lab sensitivity survives intact after deployment.
PORT and Australia are answering different operational questions
| Issue | Portugal PORT trial | Australian GP pragmatic trial | Procurement reading |
|---|---|---|---|
| Setting | Primary-care screening with MONA-GLC; 671 screened participants.[1] | Routine general-practice screening with automated retinal photography and AI.[2] | Both are prospective, but the Australian study more directly exposes routine workflow friction. |
| Sensitivity | 78%.[1] | 65% in real-world use, down from 95.6% in lab settings.[2] | Sensitivity should be priced and governed as deployed performance, not as dataset performance. |
| Specificity and referral burden | 95% specificity; 10% referred by AI versus 18% by expert panel.[1] | Follow-up confirmation remained low among those who attended specialist review.[2] | Referral reduction is valuable, but only if missed cases and nonattendance are monitored. |
| Image analyzability | Key summary figures emphasize referral and diagnostic performance.[1] | Only 66.9% had analyzable images; ungradable-image patients were older and more likely to have cataracts.[2] | Ungradable images need an explicit off-ramp, not quiet exclusion from performance reporting. |
| Economic signal | ICER of €1,725/QALY at 1% prevalence; cost-saving dominant at prevalence of 2% or higher.[1] | The trial highlights real-world workflow losses that can erode expected value.[2] | Cost-effectiveness claims need local modeling and operational assumptions. |
These two studies are not best understood as a contradiction. PORT shows what a well-performing AI triage layer may achieve in a structured pathway. The Australian trial shows what happens when the same category of intervention meets everyday primary-care constraints: fewer analyzable images, lower sensitivity, uncertain population fit, and incomplete referral completion.
Human-only referral is not the safe default
It would be too easy to compare AI with an imagined perfect clinician. Real referral pathways do not work that way. Human screening misses glaucoma, and human referral queues can be both overinclusive and unsafe. A review on undetected glaucoma in Australia reported that more than 70% of glaucoma is undiagnosed globally; in Victoria, Australia, 63% of referable glaucoma was undiagnosed, and 66% of missed cases had seen an optometrist in the prior year with cup-disc ratios greater than 0.7 that were not flagged.[4]
That is the strongest argument for standardization. AI does not need to be perfect to be useful if it reliably applies a threshold, documents the reason for triage, and creates a trackable referral decision. But it must be inserted into a system that knows when the AI did not actually assess the patient.
Human-AI collaboration needs governance, not just explanations
The Gomez et al. study is useful for governance, although it should not be treated as direct evidence of real-world screening performance. In a controlled case subset, 87 optometrists achieved 60% referral accuracy with AI support compared with 51% without AI, while the AI alone achieved 80% accuracy. The human-AI team did not catch up with the standalone AI.[3]
The uncomfortable finding is that explanations did not solve the collaboration problem. SHAP and scorecard-style explanations did not close the performance gap, and feature-importance explanations increased overreliance on incorrect AI to 92%.[3] That result does not mean explanations are useless. It means explainability should not be confused with safe oversight.
For procurement, this argues against a superficial “AI plus clinician equals safety” assumption. Oversight has to specify what the clinician is expected to check: image quality, clinical mismatch, prior glaucoma history, symptoms, risk factors, and whether an ungradable or borderline result needs human review. If the human reviewer is only asked to accept or reject a confident-looking score, the pathway has not gained much resilience.
What a safer AI glaucoma referral pathway should require

A procurement decision should start by defining the AI’s role as triage. The output should route patients to specialist referral, routine monitoring, repeat imaging, or human review. It should not be purchased or messaged as a definitive glaucoma diagnosis unless the evidence base and local governance are much stronger than the current prospective literature supports.
- Require image-quality reporting as a first-class metric. Report how many patients were screened, how many images were gradable, how many were excluded, and what happened to excluded patients.
- Create a mandatory fallback route for ungradable images. Older patients and patients with cataracts should not disappear into a technical-failure category.
- Validate performance in the local population before scaling. Training-set population fit, camera type, operator skill, and disease prevalence can all affect deployed performance.
- Monitor referral completion, not just referral generation. A screening program has not completed its job when a recommendation is printed.
- Audit by subgroup. Age, cataract status, image gradability, ethnicity where appropriate and legally permissible, and prior eyecare contact should be visible in safety reporting.
- Train clinicians on failure modes, not only interface use. The important question is not whether a user can read the AI score, but whether they know when not to trust it.
These requirements are not administrative clutter. They are the difference between an AI tool that reduces noise in the referral queue and one that improves the denominator by excluding the hardest patients.
ClinicalMind scorecard
| Criterion | Appraisal |
|---|---|
| Prospective evidence | Moderately strong. PORT and the Australian GP trial move the field beyond retrospective validation, but they show different performance under different workflow conditions.[1][2] |
| Referral reduction | Promising. PORT showed AI referral of 10% versus 18% by expert panel, with high specificity.[1] |
| Sensitivity in deployment | Caution. Reported sensitivity was 78% in PORT and 65% in the Australian pragmatic trial, with the Australian result far below prior lab performance.[1][2] |
| Image-quality and denominator risk | High concern. In Australia, only 66.9% had analyzable images, and excluded patients were older and more likely to have cataracts.[2] |
| Human-AI collaboration | Unsettled. Controlled evidence suggests AI support can improve optometrist accuracy, but explanations did not eliminate performance gaps and may increase overreliance when AI is wrong.[3] |
| Cost-effectiveness | Plausible but local. PORT’s economic findings are encouraging, but European cost assumptions should not be imported directly into US procurement.[1] |
| Appropriate role | Standardized triage with human oversight, image-quality safeguards, fallback routes, and outcome monitoring. |
| Autonomous diagnosis verdict | Not supported by the current prospective evidence. |
AI glaucoma screening can be worth buying when the purchase is really a pathway redesign: community image capture, AI triage, human review where the model is weak, explicit management of ungradable images, and closed-loop referral tracking. Bought as a diagnostic shortcut, the same technology inherits the old referral pathway’s failures and adds a new one: confidence in patients the system never actually assessed.
References
- Artificial intelligence-based glaucoma screening in primary care: a cross-sectional study and economic viability analysis — The Lancet Primary Care, 2026.
- Prospective pragmatic trial of automated retinal photography and AI glaucoma screening in Australian primary care — npj Digital Medicine, 2025.
- The explainable AI dilemma under knowledge imbalance in specialist AI for glaucoma referrals in primary care — npj Digital Medicine, 2025.
- Diagnosing glaucoma in primary eye care and the role of Artificial Intelligence applications for reducing the prevalence of undetected glaucoma in Australia — PMC, Jan et al., Eye, 2024.