Explainable 3D Convolutional Neural Networks Spatiotemporal Learning for Human Handshake Interaction Recognition
Authors
Associate Professor & Head (Retired)Department of Computer Science (Aided)Sri Ramakrishna Mision Vidyalaya College of Arts and Science Coimbatore – 641 020 .Tamilnadu, INDIA. (India)
Professor Department of Computer Science Karpagam Academy of Higher Education (KAHE) Coimbatore-641021. Tamilnadu, INDIA. (India)
Article Information
DOI: 10.51583/IJLTEMAS.2026.150700059
Subject Category: Computer Science
Volume/Issue: 15/7 | Page No: 722-732
Publication Timeline
Submitted: 2026-07-27
Accepted: 2026-08-01
Published: 2026-08-12
Abstract
Human Activity Recognition (HAR) has gained significant attention in computer vision due to its wide range of applications in surveillance, social behaviour analysis, and human–computer interaction. Among various human-to-human interactions, handshake recognition is particularly important as it represents social intention and cooperative behaviour. This study presents an efficient and interpretable deep learning framework for automatic handshake recognition from video sequences. The proposed approach employs a pretrained 3D Convolutional Neural Network (3D CNN) to directly learn spatiotemporal features, enabling effective modelling of both motion dynamics and spatial relationships between interacting individuals. The experiments are conducted using two dataset namely UT-Interaction Human Interaction Dataset and SBU Kinect Interaction dataset, focusing exclusively on the handshake interaction as the target class. The dataset provides accurate ground-truth annotations, including temporal intervals and bounding boxes, which support precise localization and reliable recognition of handshake actions. Each dataset is split into 80% for training, 10% for validation, and 10% for testing to ensure robust performance evaluation. The experimental results demonstrated that the proposed 3D CNN-based framework achieved a handshake recognition accuracy of 98.92% on the UT-Interaction dataset, representing performance improvements of 9.72%, 6.32%, and 7.12% compared to CNN, BiLSTM, and RNN models, respectively.
Keywords
Explainable AI, 3D Convolutional Neural Network (3D CNN), Spatiotemporal Learning, Human Activity Recognition (HAR), Handshake Interaction Recognition
Downloads
References
1. K. Yaseen, O.-J. Kwon, J. Kim, S. Jamil, J. Lee, and F. Ullah, Next-Gen Dynamic Hand Gesture Recognition: MediaPipe, Inception-v3 and LSTM-Based Enhanced Deep Learning Model, vol. 12, pp. 117233–117245, 2024, doi: 10.1109/ACCESS.2024.3432197. [Google Scholar] [Crossref]
2. L. I. B. López, J. A. V. Paredes, and R. M. Delgado, CNN-LSTM and Post-Processing for EMG-Based Hand Gesture Recognition, Intelligent Systems with Applications, vol. 22, pp. 1–12 2024. doi: 10.1016/j.iswa.2024.200352. [Google Scholar] [Crossref]
3. Y. Song, M. Liu, F. Wang, J. Zhu, A. Hu, and N. Sun, Gesture Recognition Based on CNN-BiLSTM for Wearable Wrist Sensors, IEEE Sensors Journal, vol. 24, no. 4, pp. 3561–3572, Feb. 2024, doi: 10.1109/JSEN.2023.3349214. [Google Scholar] [Crossref]
4. R. R. Selvaraju, M. Cogswell, A. Das, R. Vedantam, D. Parikh, and D. Batra, Grad-CAM: Visual Explanations from Deep Networks via Gradient-Based Localization, IEEE Transactions on Computer Vision and Pattern Recognition, vol. 128, no. 2, pp. 336–359, 2023. [Google Scholar] [Crossref]
5. S. A. Mahmoudi, M. R. Keyvanpour, and A. Ghorbani, Explainable Deep Learning for Human Action Recognition: A Survey and Experimental Analysis, vol. 11, pp. 124901–124920, 2023, doi: 10.1109/ACCESS.2023.3324186. [Google Scholar] [Crossref]
6. Haroon, Umair and Ullah, Amin and Hussain, Tanveer and Ullah, Waseem and Sajjad, [Google Scholar] [Crossref]
7. Muhammad and Muhammad, Khan and Lee, Mi Young and Baik, Sung Wook, "A Multi-Stream Sequence Learning Framework for Human Interaction Recognition," in IEEE Transactions on Human-Machine Systems, vol. 52, no. 3, pp. 435-444, June 2022, doi: 10.1109/THMS.2021.3138708. [Google Scholar] [Crossref]
8. Shah M, Nawaz T, Nawaz R, Rashid N, Ali MO (2025) InterAcT: A generic keypoints- [Google Scholar] [Crossref]
9. based lightweight transformer model for recognition of human solo actions and interactions in aerial videos. PLoS One 20(5): e0323314. https://doi.org/10.1371/journal.pone.0323314 [Google Scholar] [Crossref]
10. Hoangcong Le, Cheng-Kai Lu, A low-latency deep learning approach for human action [Google Scholar] [Crossref]
11. recognition in medical internet of things applications, Computers and Electrical Engineering, Volume 132, 2026, 111005, ISSN 0045-7906, https://doi.org/10.1016/j.compeleceng.2026.111005. [Google Scholar] [Crossref]
12. Liu M, Li W, He B, Wang C, Qu L. Human Action Recognition Based on 3D Convolution and Multi-Attention Transformer. Applied Sciences. 2025; 15(5):2695. https://doi.org/10.3390/app15052695 [Google Scholar] [Crossref]
13. Jin L, Fan R, Han X and Cui X (2025) Convolutional spatio-temporal sequential [Google Scholar] [Crossref]
14. inference model for human interaction behavior recognition. Front. Comput. Sci. 7:1576775. doi: 10.3389/fcomp.2025.1576775. [Google Scholar] [Crossref]