Inside a new foundation model for drug discovery: how it was built and how it works
Trained on 140 million images, an AI foundation model developed by insitro and INSIGHT at Moorfields Eye Hospital has adopted novel technical approaches to boost precision in identifying disease biomarkers for drug discovery.
Foundation models are powerful AI models trained on large-scale, diverse datasets and capable of adapting to a range of complex downstream tasks. Trained on INSIGHT's own vast collection of optical coherence tomography (OCT) retinal images, the new 'EYENOv3' model was first deployed to investigate the progression of patients with age-related macular degeneration (AMD) - a leading cause of vision loss globally.
INSIGHT data scientists and AI researchers Eden Ruffell and Dominic Williamson explain how the model was built, and how it could accelerate AMD drug discovery to benefit patients.
Q. What makes this model different from other foundation models trained on OCTs?
OCTs are three dimensional scans of the retina. Most existing models are trained using only the middle slice, or “b-scan”, taken from the centre of an OCT, discarding the rest of the data. That's true of RETFound and most other published foundation models. Other models go the other way and use the full three-dimensional OCT volume. However, these models can often be restrictive in requiring the data to have a certain fixed size, and are more expensive to train.
Our model takes a new approach. It trains on individual b-scans drawn from across the the whole OCT, not just the centre. In this way, the model can still learn from all the data in the scan, developing a better overall understanding of the retina. When adapting the model to a new task, such as predicting disease progression, it feeds each slice of the OCT through separately and combines the results using a technique called attention pooling, working out which scans matter most for the prediction. That means we can drop poor-quality slices without discarding the entire volume, and we're not locked into a fixed number of b-scans per OCT.
Q. Why is the fuller volumetric picture significant for insitro's drug discovery work?
insitro's initial use case for the model is for investigating the progression of AMD, a complex, varied disease with manifestations across the retina. A model that only sees the central slice of an OCT is throwing away most of the available information. Using the 'EYENOv3' model and utilising b-scans across the full volume supports the identification of novel AMD biomarkers that other models may miss.
Q. The model builds on Meta's DINOv3 framework. What is that in simple terms?
DINOv3 is Meta's most recent open-source computer vision system, released August 2025 and trained on around 1.6 billion natural images, such as photos of people, animals, cars and everyday objects. It uses an architecture called a vision transformer, and a training method called “self-supervised contrastive learning”. In practice, the model is shown two versions of the same picture, say a dog and a cropped or filtered version of that dog, and it learns to represent both in a mathematically similar way. To do well at this, the model has to develop a proper understand of the composition of the images and what makes any two images similar or different to each other. Repeating this process across billions of images allows the model to build a general understanding of what objects are.
For developing 'EYENOv3', we took this DINOv3 model and further adapted it to the OCT scans at INSIGHT. To do this, we used anonymised unlabelled data from around 260,000 patients, totalling 2.6 million OCT volumes and 140 million individual B-scans. After training, we were then able to adapt the model for a range of downstream clinical applications, starting with mapping AMD progression.
Q. How was that model then adapted for AMD-specific work?
Fine-tuning for AMD used a dataset of around 60,000 OCTs from more than 20,000 patients. We derived disease labels for each OCT in a patient's history by cross-referencing predictions from our in-house disease classification model with anonymised treatment records. Using these two independent sources, the team could assign each OCT to an AMD disease stage and use that staging to fine-tune the foundation model. In this way we were able to train the model to specialise in understanding how a patient's disease may progress over the next few years.
Extracting insights from this predictive model can help in identifying new biomarkers of AMD in its earlier stages, feeding the model's outputs into insitro's genetic discovery work. Another use case is in clinical trials: identifying which participants are at risk of progressing quickly to late-stage AMD could aid in trial recruitment and monitoring.
Q. Could the same approach to AI model development work outside ophthalmology?
We think so. We have shown the benefits of adapting a state-of-the-art 2D foundation model, like DINOv3, to build a robust OCT foundation model capable of adapting to many tasks. Researchers using other 3D medical data, such as MRI, could plausibly benefit from the same combination: a strong general-purpose base model adapted with slices taken from across other 3D domain-specific data.
Q. When can we expect to see published results?
A paper on how we built the foundation model will be published in the coming months. The AMD-specific work will be detailed in a separate paper to follow, explaining how we built the labelled dataset and demonstrating the effectiveness of the approach we took.

Eden Ruffell is an INSIGHT Data Scientist and AI Researcher, and Honorary Research Fellow at Moorfields Eye Hospital. She is also a doctoral researcher at UCL, researching the development of OCT foundation models, supervised by Professor Pearse Keane and Professor Daniel Alexander.

Dominic Williamson is an INSIGHT Data Scientist and AI Researcher, and Honorary Research Fellow at Moorfields Eye Hospital. He is also a doctoral researcher at UCL, using AI to research the manifestation of systemic and ocular diseases in the retina, supervised by Professor Pearse Keane and Professor Daniel Alexander.
About insitro
Founded and led by AI pioneer Daphne Koller, insitro is a machine learning-enabled drug discovery and development company creating a new approach for target and drug discovery. insitro is uncovering genetic targets and new therapeutic hypotheses by integrating multimodal data from human cohorts and cellular models with the power of AI and machine learning to increase the therapeutic probability of success. These insights provide the starting point for discovering new molecules, which are either built with in-house, AI-enabled drug discovery platforms or with partners that extend insitro’s impact. For more about insitro: https://www.insitro.com/



