AI applications in retinal imaging range from FDA-approved diabetic retinopathy screening systems (supported evidence) to experimental disease prediction models (preliminary evidence). Real-world performance, external validation, and clinical integration remain critical challenges across most applications.
The scope of AI in ophthalmology
Artificial intelligence, particularly deep learning, has been applied to fundus photography, OCT, and OCTA for tasks including disease detection, severity grading, progression prediction, and discovering associations with systemic conditions. Some applications have reached clinical deployment; most remain investigational.
Diabetic retinopathy screening
Multiple AI systems for automated DR screening have received FDA clearance, including autonomous diagnostic systems. These show high sensitivity and specificity in controlled studies. However, real-world deployment faces challenges: ungradable image rates, performance variation across populations and camera types, and workflow integration questions. Evidence level: Supported for specific approved systems.
AMD and OCT interpretation
AI models have been developed for detecting and quantifying fluid on OCT, measuring geographic atrophy, classifying choroidal neovascularization, and predicting AMD progression. Research performance is strong, but prospective clinical validation is limited. Some systems are entering clinical use as decision support tools, not autonomous diagnostics. Evidence level: Supported for detection tasks, Emerging for progression prediction.
3D volumetric OCT analysis
Newer approaches analyze entire OCT volumes rather than selected B-scans, potentially capturing spatial information clinicians use. This could improve diagnostic accuracy and support more complex clinical decisions. However, computational requirements are high, training data limited, and external validation across devices incomplete. Evidence level: Emerging.
Foundation models for ophthalmology
Large AI models pre-trained on diverse retinal images, then fine-tuned for specific tasks, represent a recent development. These models may generalize better than task-specific networks and require less labeled data for new applications. Published benchmarks show promise, but clinical deployment and prospective validation are very early stage. Evidence level: Emerging.
Glaucoma detection and monitoring
AI applied to OCT nerve fiber layer analysis, optic disc photography, and visual fields shows research promise for earlier detection and progression monitoring. Integration with existing clinical workflow and demonstrating improved patient outcomes are ongoing challenges. Evidence level: Emerging to Supported depending on specific application.
Critical limitations across applications
Most AI systems face common challenges: they may not generalize well to populations, devices, or image quality conditions different from training data. Explainability remains difficult—understanding why an AI made a specific prediction is not always possible. Regulatory pathways for updating deployed models are unclear. Long-term clinical outcome studies are rare. Real-world ungradable rates often exceed research settings.
What deployment requires
Moving from research benchmarks to clinical practice requires prospective validation, integration into existing workflows, clear liability frameworks, continuous performance monitoring, and evidence that AI improves care—not just matches human performance on retrospective datasets. Regulatory approval is necessary but not sufficient for clinical adoption.
Evidence should be inspectable.
This article is part of the earlier V1 library. We are progressively upgrading each piece with primary literature, structured references and explicit limitations.
Read our editorial standard →