AI Model Achieves 90% Accuracy Detecting Glaucoma Risk in UK Study
Machine learning system outperformed trained human graders analyzing retinal photographs from more than 6,000 participants.
AI system demonstrates superior glaucoma detection in population screening
A machine learning model has demonstrated significantly higher accuracy than trained human graders in detecting glaucoma risk from retinal photographs, according to findings from a large-scale population study in the United Kingdom.
The AI system was tested against data from 6,304 participants in the EPIC-Norfolk Eye Study, a long-running population health research project. The model achieved an area under the receiver operating characteristic curve (AUROC) between 88% and 90% when predicting specialist-confirmed glaucoma cases from retinal images.
This performance metric indicates the model's ability to correctly distinguish between participants with and without glaucoma risk substantially exceeded that of human graders who performed the same task. The study represents a real-world validation of AI diagnostic capabilities in ophthalmology, testing the technology against a diverse population sample rather than curated clinical datasets.
Why it matters
Glaucoma remains a leading cause of irreversible blindness worldwide, yet early detection through screening can prevent vision loss. The superior accuracy demonstrated by this AI model addresses a critical bottleneck in population-level screening programs: the shortage of trained specialists to review large volumes of retinal images. If deployed at scale, such systems could enable more comprehensive screening coverage while reducing the burden on human graders and potentially catching cases that might otherwise be missed.
Implications for vision screening programs
The researchers behind the study indicated that these results strengthen the argument for integrating AI-assisted screening into public health vision programs. Current screening approaches often rely on manual review of retinal photographs by trained graders, a process that is time-intensive and subject to human variability and fatigue.
An AI system capable of matching or exceeding human performance could process far larger volumes of images while maintaining consistent diagnostic standards. This scalability becomes particularly important for aging populations where glaucoma prevalence increases and healthcare systems face growing demand for screening services.
The EPIC-Norfolk Eye Study provided an especially rigorous testing environment because it drew from a general population sample rather than patients already suspected of having eye disease. This real-world distribution of healthy and at-risk individuals more closely mirrors the conditions AI systems would encounter in actual screening deployments.
Technical performance and validation
The AUROC metric used to evaluate the model represents a standard measure in diagnostic testing. Scores between 88% and 90% indicate strong discriminatory ability, though not perfect classification. For context, an AUROC of 50% would represent performance no better than random chance, while 100% would indicate flawless classification.
The study validated the AI model against specialist-confirmed diagnoses, ensuring that the ground truth used for comparison represented expert clinical judgment rather than potentially error-prone initial assessments.
These findings were first reported by Ophthalmology Times.
This is an original analysis by the Omega editorial team. Source reporting: AI Watch.
Want systems like this working for your business?
Book a Call

