Explainable Deep Learning for Intelligent Plant Disease Detection
Authors
Dr. Pallavi Sharma
Assistant Professor, School of Engineering, Design and Automation- E, Department of ECE, GNA University, Phagwara, Punjab, India (IN)
Ngah Hesly Kilofonyuy
Undergraduate Student, School of Engineering, Design and Automation- E, Department of ECE, GNA University, Phagwara, Punjab, India (IN)
Article Information
DOI: 10.51583/IJLTEMAS.2026.150500184
Subject Category: Deep Learning
Volume/Issue: 15/5 | Page No: 2297-2314
Publication Timeline
Submitted: 2026-06-12
Published: 2026-06-12
Abstract
The world suffers from 10–40% loss in crop yields each year because of plant disease. This threat is serious and growing; it threatens food security, rural livelihoods, and agricultural economies. Advances being made through deep learning, computer vision, and mobile technology have presented a unique opportunity to use leaf images to automatically recognize plant disease. Published classification accuracies on benchmark datasets now exceed 97%, which is an important achievement but achieving high accuracy on a benchmark alone does not indicate that traditional methods will work when deployed in the real world: all four stakeholders (i.e., farmers, agronomists, regulatory authorities, and extension agents) must therefore have the ability to understand, and interpret the output of automatically recognized plant diseases in a way that enhances human expertise rather than replacing it. In this chapter, we provide a compendium of technical deep learning architectures and methods related to Explainable Artificial Intelligence (XAI) for plant disease detection, including convolutional networks, residual architectures, dense architectures, transformer networks, and hybrid models. We also systematically evaluate the explainability methods used in both post-hoc and intrinsic explanation and evaluate the applicability of these methods across a variety of imaging modalities used in agriculture, including RGB, multispectral, and hyperspectral. This chapter characterizes major benchmark datasets; discusses major challenges to their deployment, including class imbalance, domain shift, model size reduction, and human–AI trust calibration; then ends with potential new directions for research in areas such as foundation models (FM), causal interpretable models (Explanations), federated learning, and continual learning to build resilience for each evolving pathogen landscape.
Keywords
Explainable Artificial Intelligence; Plant Disease Detection; LIME, Grad-CAM; SHAP; Federated Learning; Precision Agriculture; Hyperspectral Imaging.
Downloads
References
1. Mohanty, S. P., Hughes, D. P., & Salathé, M. (2016). Using deep learning for image-based plant disease detection. Frontiers in Plant Science, 7, Article 1419. https://doi.org/10.3389/fpls.2016.01419 [Google Scholar] [Crossref]
2. Simonyan, K., & Zisserman, A. (2015). Very deep convolutional networks for large-scale image recognition. International Conference on Learning Representations (ICLR). [Google Scholar] [Crossref]
3. He, K., Zhang, X., Ren, S., & Sun, J. (2016). Deep residual learning for image recognition. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (pp. 770–778). https://doi.org/10.1109/CVPR.2016.90 [Google Scholar] [Crossref]
4. Huang, G., Liu, Z., Van Der Maaten, L., & Weinberger, K. Q. (2017). Densely connected convolutional networks. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (pp. 4700–4708). https://doi.org/10.1109/CVPR.2017.243 [Google Scholar] [Crossref]
5. Tan, M., & Le, Q. V. (2019). EfficientNet: Rethinking model scaling for convolutional neural networks. In Proceedings of the 36th International Conference on Machine Learning (pp. 6105–6114). [Google Scholar] [Crossref]
6. Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., et al. (2020). An image is worth 16×16 words: Transformers for image recognition at scale. arXiv. https://arxiv.org/abs/2010.11929 [Google Scholar] [Crossref]
7. Liu, Z., Lin, Y., Cao, Y., Hu, H., Wei, Y., Zhang, Z., et al. (2021). Swin Transformer: Hierarchical vision transformer using shifted windows. In Proceedings of the IEEE/CVF International Conference on Computer Vision (pp. 10012–10022). https://doi.org/10.1109/ICCV48922.2021.00986 [Google Scholar] [Crossref]
8. Doshi-Velez, F., & Kim, B. (2017). Towards a rigorous science of interpretable machine learning. arXiv. https://arxiv.org/abs/1702.08608 [Google Scholar] [Crossref]
9. Selvaraju, R. R., Cogswell, M., Das, A., Vedantam, R., Parikh, D., & Batra, D. (2017). Grad-CAM: Visual explanations from deep networks via gradient-based localization. In Proceedings of the IEEE International Conference on Computer Vision (pp. 618–626). https://doi.org/10.1109/ICCV.2017.74 [Google Scholar] [Crossref]
10. Achanta, R., Shaji, A., Smith, K., Lucchi, A., Fua, P., & Süsstrunk, S. (2012). SLIC superpixels compared to state-of-the-art superpixel methods. IEEE Transactions on Pattern Analysis and Machine Intelligence, 34(11), 2274–2282. https://doi.org/10.1109/TPAMI.2012.120 [Google Scholar] [Crossref]
11. Ribeiro, M. T., Singh, S., & Guestrin, C. (2016). “Why should I trust you?”: Explaining the predictions of any classifier. In Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining (pp. 1135–1144). https://doi.org/10.1145/2939672.2939778 [Google Scholar] [Crossref]
12. Lundberg, S. M., & Lee, S. I. (2017). A unified approach to interpreting model predictions. In Advances in Neural Information Processing Systems (Vol. 30, pp. 4765–4774). [Google Scholar] [Crossref]
13. Abnar, S., & Zuidema, W. (2020). Quantifying attention flow in transformers. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics (pp. 4190–4197). https://doi.org/10.18653/v1/2020.acl-main.385 [Google Scholar] [Crossref]
14. Woo, S., Park, J., Lee, J. Y., & Kweon, I. S. (2018). CBAM: Convolutional block attention module. In Proceedings of the European Conference on Computer Vision (pp. 3–19). https://doi.org/10.1007/978-3-030-01234-2_1 [Google Scholar] [Crossref]
15. Samek, W., Montavon, G., Lapuschkin, S., Anders, C. J., & Müller, K. R. (2021). Explaining deep neural networks and beyond: A review of methods and applications. Proceedings of the IEEE, 109(3), 247–278. https://doi.org/10.1109/JPROC.2021.3060483 [Google Scholar] [Crossref]
16. Hughes, D. P., & Salathé, M. (2015). An open access repository of images on plant health to enable the development of mobile disease diagnostics. arXiv. https://arxiv.org/abs/1511.08060 [Google Scholar] [Crossref]
17. Singh, D., Jain, N., Jain, P., Kayal, P., Kumawat, S., & Batra, N. (2020). PlantDoc: A dataset for visual plant disease detection. In Proceedings of the 7th ACM IKDD CoDS and 25th COMAD (pp. 249–253). https://doi.org/10.1145/3371158.3371196 [Google Scholar] [Crossref]
18. Lin, T. Y., Goyal, P., Girshick, R., He, K., & Dollár, P. (2017). Focal loss for dense object detection. In Proceedings of the IEEE International Conference on Computer Vision (pp. 2980–2988). https://doi.org/10.1109/ICCV.2017.324 [Google Scholar] [Crossref]
19. Radford, A., Kim, J. W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., et al. (2021). Learning transferable visual models from natural language supervision. In Proceedings of the 38th International Conference on Machine Learning (pp. 8748–8763). [Google Scholar] [Crossref]
20. Chen, T., Kornblith, S., Norouzi, M., & Hinton, G. (2020). A simple framework for contrastive learning of visual representations. In Proceedings of the 37th International Conference on Machine Learning (pp. 1597–1607). [Google Scholar] [Crossref]
21. McMahan, B., Moore, E., Ramage, D., Hampson, S., & Agüera y Arcas, B. (2017). Communication-efficient learning of deep networks from decentralized data. In Proceedings of the 20th International Conference on Artificial Intelligence and Statistics (pp. 1273–1282). [Google Scholar] [Crossref]
22. Chollet, F. (2017). Xception: Deep learning with depthwise separable convolutions. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (pp. 1251–1258). https://doi.org/10.1109/CVPR.2017.195 [Google Scholar] [Crossref]
23. Zhou, B., Khosla, A., Lapedriza, A., Oliva, A., & Torralba, A. (2016). Learning deep features for discriminative localization. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (pp. 2921–2929). https://doi.org/10.1109/CVPR.2016.319 [Google Scholar] [Crossref]
Metrics
Views & Downloads
Similar Articles
- Block-Based Programming for Education: A Comprehensive Analysis of Visual Programming Environments in K-12 Learning
- Management of Academic Libraries and Client Satisfaction Towards Digital Utilization: Basis for Monitoring Library Operations in SOCCSKSARGEN Region.
- Revenue Leakages in TPA Insurance Claims and Corporate Claims: An Institutional Overview of Aster Prime Hospital, Hyderabad
- Technology and Innovation in Hospitality and Tourism: A Management Perspective
- Geospatial Distribution of Tarok Sacred Grove of Langtang North and Langtang South Local Government Areas