00
Days
00
Hrs
00
Min
00
Sec
Submit Your Paper

A Comparative Study of CNN and Vision Transformer Models in Alzheimer's Detection

Authors

D Stephen

Research Scholar, Department of Computer Science and Engineering, Saveetha School of Engineering, Saveetha Institute of Medical and Technical Sciences (SIMATS), Chennai, India (India)

Dr. R Samuel Rajesh Babu

Associate Professor, Department of Electronics and Communication Engineering, Saveetha School of Engineering, Saveetha Institute of Medical and Technical Sciences (SIMATS), Chennai, India (India)

Article Information

DOI: 10.51583/IJLTEMAS.2026.150800141

Subject Category: Machine Learning

Volume/Issue: 15/8 | Page No: 1952-1960

Publication Timeline

Submitted: 2026-09-03

Accepted: 2026-09-08

Published: 2026-09-26

Abstract

In this work, a novel deep learning framework the Hierarchical Attention Network (HierAttNet) is proposed to replace the previously used Vision Transformer (ViT-B/16) and all CNN baselines in the literature to stage Alzheimer's disease (AD) from MRI brain images. The ConvNeXt-Small feature extractor and the two-scale cross-attention module alternately attend to the local patch level and the region level in a hierarchical fashion, which is not reported in the AD neuroimaging literature. We compare HierAttNet with six additional state-of-the-art algorithms, which were not presented in the original paper: EfficientNetV2-S, Swin-T, MobileViT-S, CrossViT-15, DeiT-S and ConvNeXt-S. The following experiments are performed on the Kaggle Alzheimer's MRI dataset (6,400+ images, 4 stages of Alzheimer's). HierAttNet obtains 99.91% accuracy and F1=0.9993 outperforming all the evaluated models and even the previous state-of-the-art ViT-B/16 (99.87%). The performance of all models consistently dropped with data augmentation. In this work, we introduce HierAttNet as a new state-of-the-art in Alzheimer's MRI classification, and a stringent benchmark for subsequent multimodal diagnostics studies.

Keywords

Index Terms Alzheimer's disease, MRI classification, hierarchical attention, vision transformers, deep learning, HierAttNet, convolutional neural networks.

Downloads

References

1. M. Tan and Q. V. Le, "EfficientNetV2: Smaller Models and Faster Training," in Proc. ICML, 2021. [Google Scholar] [Crossref]

2. Z. Liu et al., "A ConvNet for the 2020s," in Proc. IEEE CVPR, 2022. [Google Scholar] [Crossref]

3. Z. Liu et al., "Swin Transformer: Hierarchical Vision Transformer using Shifted Windows," in Proc. IEEE ICCV, 2021. [Google Scholar] [Crossref]

4. S. Mehta and M. Rastegari, "MobileViT: Light-weight, General-purpose, and Mobile-friendly Vision Transformer," arXiv:2110.02178, 2021. [Google Scholar] [Crossref]

5. C.-F. R. Chen et al., "CrossViT: Cross-Attention Multi-Scale Vision Transformer for Image Classification," in Proc. IEEE ICCV, 2021. [Google Scholar] [Crossref]

6. H. Touvron et al., "Training Data-Efficient Image Transformers & Distillation through Attention," in Proc. ICML, 2021. [Google Scholar] [Crossref]

7. Dosovitskiy et al., "An Image is Worth 16×16 Words: Transformers for Image Recognition at Scale," in Proc. ICLR, 2021. [Google Scholar] [Crossref]

8. Radosavovic et al., "Designing Network Design Spaces," in Proc. IEEE CVPR, 2020. [Google Scholar] [Crossref]

9. J. Hu et al., "Squeeze-and-Excitation Networks," IEEE CVPR, 2018. [Google Scholar] [Crossref]

10. G. Huang et al., "Densely Connected Convolutional Networks," in Proc. IEEE CVPR, 2017. [Google Scholar] [Crossref]

11. P. LaMontagne et al., "OASIS-3: Longitudinal Neuroimaging, Clinical and Cognitive Dataset," medRxiv, 2019. [Google Scholar] [Crossref]

12. H. Touvron et al., "ResMLP: Feedforward networks for image classification with data-efficient training," IEEE TPAMI, 2022. [Google Scholar] [Crossref]

Metrics

Views & Downloads

Similar Articles

© 2026 IJLTEMAS · RSIS International. All rights reserved. ISSN 2278-2540.