Explainable Vision Transformer for Diabetic Retinopathy Detection
DOI:
https://doi.org/10.65890/dmp-lncse.ICICCS26.218Keywords:
Diabetic Retinopathy, Vision Transformer, Explainable AI, Medical Image Analysis, Deep Learning, Attention Mechanism, Fundus ImagingAbstract
Diabetic Retinopathy (DR) is a major cause of vision loss all over the world, and if diagnosed early, it can be prevented. In this paper, we propose an automated retinal image classifier using an Explainable Vision Transformer (ViT) model to classify Diabetic Retinopathy (DR) on the APTOS 2019 retinal image dataset. It leverages the transformer's attention mechanism to focus on clinically relevant regions in retinal fundus images, demonstrating high interpretability and strong predictive performance. Pre-processing techniques to enhance model generalisation include image resizing, image normalisation, and image augmentation. ViT models have been trained to achieve about 95 per cent accuracy on the validation set for predicting five levels of DR severity. Moreover, attention maps visually illustrate regions that significantly impact the model's decision-making process and, in addition, provide interpretable outputs for medical experts. The proposed strategy highlights the opportunities afforded by transformer-based architectures for medical imaging, with their high-performance enabling safety and reliability in DR screening.
Downloads
Published
Conference Proceedings Volume
Section
License
Copyright (c) 2026 DMPedia Lecture Notes in Computer Science & Engineering

This work is licensed under a Creative Commons Attribution 4.0 International License.