Inspiration

Diabetic Retinopathy (DR) is a major complication of diabetes that can progressively damage vision. We were inspired by the idea that AI could assist in detecting DR from retinal fundus images while also considering relevant patient information.

Instead of relying only on visual features from an image, we wanted to explore a multimodal approach that combines retinal images with clinical attributes. This led us to develop a Vision Transformer (ViT)-based framework that can first determine whether DR is present and, for DR-positive cases, identify its severity.

What it does

Our project, Multimodal ViT Framework for DR Detection & Severity Grading, analyzes retinal fundus images together with selected clinical information.

The system follows a two-stage hierarchical approach:

Stage 1 — DR Detection: Determines whether the retinal image indicates No DR or DR. Stage 2 — Severity Grading: If DR is detected, the system further classifies it as: Mild Moderate Severe Proliferative DR

Together, these stages preserve the standard five-level grading scheme:

$$ Y_{DR} \in {0,1,2,3,4} $$

where 0 represents No DR and 1–4 represent increasing DR severity.

The framework combines retinal image features extracted using ViT-B/16 with clinical information processed through a clinical encoder, followed by cross-modal attention to create a multimodal representation.

How we built it

We built the framework using the BRSET (Brazilian Multilabel Ophthalmological Dataset), combining retinal fundus images with selected patient-level clinical attributes.

The image branch uses ViT-B/16 to extract visual representations from retinal fundus images. The clinical branch processes seven selected attributes:

Patient age Diabetes duration Insulin Patient sex Exam eye Diabetes Comorbidities

The clinical features are transformed into a compact representation, which is then combined with retinal image features through cross-modal attention.

The resulting multimodal representation is passed through the hierarchical prediction process:

$$ (I,C) \rightarrow \text{ViT}(I),\text{MLP}(C) \rightarrow \text{Cross-Modal Attention} \rightarrow Z \rightarrow \text{Stage 1} \rightarrow \text{Stage 2} $$

Stage 2 is conditionally executed only when Stage 1 identifies DR.

We also incorporated Grad-CAM-based explainability to help visualize the retinal regions contributing to model predictions.

Challenges we ran into

One of the main challenges was combining two different types of information—high-dimensional retinal images and structured clinical data—into a meaningful representation.

We also had to design the prediction process carefully so that the system could distinguish between detecting the presence of DR and grading its severity. Handling multiple severity levels and class imbalance in the dataset added further complexity.

Another challenge was making the model more interpretable. Since medical AI systems should not simply provide a prediction without context, we explored explainability using Grad-CAM to understand which retinal regions influenced the model's decision.

Accomplishments that we're proud of

We are proud of developing a multimodal and hierarchical framework rather than relying solely on retinal images.

Our framework:

Combines retinal fundus images and clinical information. Uses ViT-B/16 for visual feature extraction. Uses cross-modal attention to connect image and clinical representations. Separates DR detection from severity grading through a two-stage approach. Supports the complete five-level ICDR grading scheme. Incorporates Grad-CAM for model explainability. Achieved a DR-positive detection AUC of 0.84 and F1-score of 0.78, with a QWK of 0.77 for severity grading.

Most importantly, the project helped us move from simply building a classifier toward thinking about how a multimodal AI system could be structured for a real medical screening scenario.

What we learned

Through this project, we learned that medical AI is not only about achieving high accuracy. Data quality, clinical relevance, class imbalance, explainability, and the way predictions are structured are equally important.

We gained practical experience with:

Vision Transformers and ViT-B/16 Multimodal learning Clinical feature preprocessing Cross-modal attention Hierarchical classification Medical image analysis Model explainability with Grad-CAM Working with real-world ophthalmological datasets

We also learned the importance of designing AI systems in a way that makes their predictions easier to understand and evaluate.

What's next for Multimodal ViT Framework for DR Detection Severity Grading

Our next step is to make the framework more robust and clinically meaningful.

We plan to explore more powerful multimodal architectures and improved interaction between retinal and clinical representations. We also aim to evaluate the framework on larger and more diverse datasets and improve performance on underrepresented severity grades.

Future work will also consider incorporating additional clinical information such as glycemic control, blood pressure, treatment history, and other relevant clinical measurements, where appropriate data is available.

For explainability, we want to compare Grad-CAM visualizations with expert-annotated retinal lesions and clinical review to better understand whether the model is focusing on medically meaningful regions.

Ultimately, our goal is to develop a more robust, explainable, and generalizable multimodal AI framework that can support diabetic retinopathy screening and severity assessment.

Built With

Share this project:

Updates

Submission history