A Lightweight Hybrid CNN-Transformer Model for Early Retinal Disease Screening in Low-Resource Clinical Settings
DOI:
https://doi.org/10.71229/aveayq30Keywords:
retinal fundus imaging,, multi-disease screening,, hybrid CNN–Transformer,, degradation-aware training,, model robustness, , low-resource healthcare,Abstract
Automated retinal screening has the potential to provide specialist-level triage to primary care settings, but state-of-the-art models trained on high-quality fundus photographs tend to degrade sharply when evaluated on handheld or smartphone-based camera-captured images, which are precisely the kind of equipment found in low-resource environments. We introduce a lightweight hybrid CNN–Transformer model which features a MobileNetV3-Small stem coupled with a strided depth-wise convolutional token embedding and two narrow transformer blocks, and which is trained with a dual head for both binary disease-risk screening and 25-label multi-disease classification. The model is trained with degradation-aware training (DAT), in which one of five simulated acquisition artefacts — defocus blur, weak illumination, JPEG compression, low sensor resolution and read noise — is applied at a randomly drawn severity, which is never equal to the test severity. Evaluated on the RFMiD dataset (three seeds), the model attains a clean screening AUC of 0.966 ± 0.001 at 2.3 M parameters and 0.12 GFLOPs, within 0.01 of far larger baselines, while retaining an AUC of 0.936 under severe simulated degradation, where EfficientNet-B0 and ViT-Tiny degrade to 0.603 and 0.777 and invert below chance on individual corruptions. Matched-architecture ablations attribute +0.180 AUC to DAT alone.
References
[2] Dong, L., He, W., Zhang, R., Ge, Z., Wang, Y. X., Zhou, J., . . . Wei, W. B. (2022). Artificial Intelligence for Screening of Multiple Retinal and Optic Nerve Diseases. JAMA Network Open, 5(5), e229960-e229960. doi:10.1001/jamanetworkopen.2022.9960 DOI: https://doi.org/10.1001/jamanetworkopen.2022.9960
[3] Elkenawy, E.-S. M., Khodadadi, N., Gaber, K. S., Khodadadi, E., Alhussan, A. A., Khafaga, D. S., & Eid, M. M. (2026). Automated retinal disease classification using deep learning and AlexNet with statistical models analysis. PLOS One, 21(1), e0338415. doi:10.1371/journal.pone.0338415 DOI: https://doi.org/10.1371/journal.pone.0338415
[4] Hendrycks, D., & Dietterich, T. (2019). Benchmarking Neural Network Robustness to Common Corruptions and Perturbations. arXiv preprint, arXiv.1903.12261. doi:10.48550/arXiv.1903.12261
[5] Howard, A., Sandler, M., Chen, B., Wang, W., Chen, L. C., Tan, M., . . . Le, Q. (2019). Searching for MobileNetV3. Paper presented at the 2019 IEEE/CVF International Conference on Computer Vision (ICCV), Seoul, Korea (South). https://doi.org/10.1109/ICCV.2019.00140 DOI: https://doi.org/10.1109/ICCV.2019.00140
[6] Lin, J., Zheng, J., & Lin, B. (2025). A review of deep learning for fundus image enhancement. Discover Computing, 28(1), 233. doi:10.1007/s10791-025-09761-5 DOI: https://doi.org/10.1007/s10791-025-09761-5
[7] Loshchilov, I., & Hutter, F. (2017). SGDR: Stochastic Gradient Descent with Warm Restarts. arXiv preprint, arXiv.1608.03983. doi:10.48550/arXiv.1608.03983
[8] Loshchilov, I., & Hutter, F. (2019). Decoupled Weight Decay Regularization. arXiv preprint, arXiv:1711.05101. doi:10.48550/arXiv.1711.05101
[9] Mehta, S., & Rastegari, M. (2022). MobileViT: Light-weight, General-purpose, and Mobile-friendly Vision Transformer. arXiv preprint, arXiv.2110.02178. doi:10.48550/arXiv.2110.02178
[10] Menaka, S. R., Muthusamy, S., Sidhu, P. K., Routray, A., Maheswari, G. U., & Bacanin, N. (2026). A novel diabetic retinopathy detection from fundus images using hybrid quantum convolutional neural network models. Scientific Reports, 16(1), 20878. doi:10.1038/s41598-026-49227-2 DOI: https://doi.org/10.1038/s41598-026-49227-2
[11] Muchuchuti, S., & Viriri, S. (2023). Retinal Disease Detection Using Deep Learning Techniques: A Comprehensive Review. Journal of Imaging, 9(4), 84. doi:10.3390/jimaging9040084 DOI: https://doi.org/10.3390/jimaging9040084
[12] Oliveira, G. C., Rosa, G. H., Pedronette, D. C. G., Papa, J. P., Kumar, H., Passos, L. A., & Kumar, D. (2024). Robust deep learning for eye fundus images: Bridging real and synthetic data for enhancing generalization. Biomedical Signal Processing and Control, 94, 106263. doi:10.1016/j.bspc.2024.106263 DOI: https://doi.org/10.1016/j.bspc.2024.106263
[13] Pachade, S., Porwal, P., Kokare, M., Deshmukh, G., Sahasrabuddhe, V., Luo, Z., . . . Mériaudeau, F. (2025). RFMiD: Retinal Image Analysis for multi-Disease Detection challenge. Medical Image Analysis, 99, 103365. doi:10.1016/j.media.2024.103365 DOI: https://doi.org/10.1016/j.media.2024.103365
[14] Panchal, S., Naik, A., Kokare, M., Pachade, S., Naigaonkar, R., Phadnis, P., & Bhange, A. (2023). Retinal Fundus Multi-Disease Image Dataset (RFMiD) 2.0: A Dataset of Frequently and Rarely Identified Diseases. Data, 8(2), 29. doi:10.3390/data8020029 DOI: https://doi.org/10.3390/data8020029
[15] Ridnik, T., Ben-Baruch, E., Zamir, N., Noy, A., Friedman, I., Protter, M., & Zelnik-Manor, L. (2021). Asymmetric Loss For Multi-Label Classification. Paper presented at the 2021 IEEE/CVF International Conference on Computer Vision (ICCV), Montreal, QC, Canada. https://doi.org/10.1109/ICCV48922.2021.00015 DOI: https://doi.org/10.1109/ICCV48922.2021.00015
[16] Sah, U. K., Chatterjee, J. M., & Sujatha, R. (2026). Multi-class eye disease classification using deep learning EfficientNetB0 fusion techniques. Scientific Reports, 16(1), 6368. doi:10.1038/s41598-026-35357-0 DOI: https://doi.org/10.1038/s41598-026-35357-0
[17] Savoy, F. M., Rao, D. P., Toh, J. K., Ong, B., Sivaraman, A., Sharma, A., & Das, T. (2024). Empowering Portable Age-Related Macular Degeneration Screening: Evaluation of a Deep Learning Algorithm for a Smartphone Fundus Camera. BMJ Open, 14(9), e081398. doi:10.1136/bmjopen-2023-081398 DOI: https://doi.org/10.1136/bmjopen-2023-081398
[18] Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., . . . Polosukhin, I. (2023). Attention Is All You Need. arXiv preprint, arXiv.1706.03762. doi:10.48550/arXiv.1706.03762
[19] Wang, X., Zhu, Y., Cui, Y., Huang, X., Guo, D., Mu, P., . . . Chen, S. (2025). Lightweight Multi-Stage Aggregation Transformer for robust medical image segmentation. Medical Image Analysis, 103, 103569. doi:10.1016/j.media.2025.103569 DOI: https://doi.org/10.1016/j.media.2025.103569
Downloads
Published
Issue
Section
License
Copyright (c) 2026 Al-Noor Journal of Engineering Management and Computer Science

This work is licensed under a Creative Commons Attribution-NonCommercial-ShareAlike 4.0 International License.





