Foundation models pre-trained on large retinal image datasets, then adapted for specific tasks, represent a paradigm shift in ophthalmic AI. Early research shows strong performance on benchmark tasks and potential for data-efficient transfer learning. Prospective clinical validation, fairness across demographics, and real-world generalization remain largely untested.
What foundation models are
In AI, a foundation model is a large neural network pre-trained on massive datasets, learning general representations that transfer to many downstream tasks. In retinal imaging, this means training on hundreds of thousands or millions of fundus photos, OCT scans, or multimodal images, then fine-tuning for specific diseases or prediction tasks.
Why this approach differs
Traditional retinal AI trains task-specific models from scratch: one model for diabetic retinopathy, another for AMD, another for glaucoma. Foundation models learn generalizable retinal features once, then adapt quickly to new tasks with less labeled data. This could accelerate development of AI for rare diseases or new applications.
Early results
Models like RETFound (published 2023) and others show strong performance when fine-tuned on tasks including disease detection, progression prediction, and oculomics associations. These models sometimes outperform or match task-specific models, especially when training data for the target task is limited. Results are promising but from retrospective test sets.
The data efficiency claim
Foundation model advocates argue they require fewer labeled examples for new tasks (few-shot or zero-shot learning). If true, this could enable AI for diseases with limited annotated data. However, whether this generalizes across real clinical populations and imaging conditions remains uncertain.
Generalization questions
Do foundation models generalize better across different OCT devices, camera types, populations, and disease severities than task-specific models? Early evidence is mixed. Benchmark improvements of 1-2% AUC may not translate to clinically meaningful differences. External validation across diverse real-world settings is limited.
Fairness and bias
If foundation models are trained on non-representative datasets, they may perpetuate or amplify existing biases. Performance across race, ethnicity, age, and disease stage must be rigorously validated. Some studies show performance gaps; others do not. Transparency about training data demographics is often lacking.
Explainability challenges
Foundation models are even larger and more complex than previous retinal AI systems, making interpretation harder. Attention maps and saliency methods provide some insight but are imperfect. Regulatory bodies and clinicians may require better explainability for high-stakes medical decisions.
What would validate the promise
Prospective clinical trials comparing foundation-model-based systems to current standards. Evidence of improved patient outcomes, not just benchmark metrics. Demonstration of fairness across demographics. Clear value proposition: faster development for rare diseases? Better generalization? More accurate diagnosis? The answers will determine whether foundation models transform retinal care or remain a research curiosity.
Evidence should be inspectable.
This article is part of the earlier V1 library. We are progressively upgrading each piece with primary literature, structured references and explicit limitations.
Read our editorial standard →