Vision Transformer Based Digital Image Forgery Detection and Localization Using Global Contextual Feature Learning
Article Sidebar
Main Article Content
Artificial intelligence has significantly improved digital image editing capabilities, making it increasingly difficult to distinguish authentic images from manipulated ones [5, 7]. This paper proposes a Vision Transformer (ViT)-based framework for digital image forgery detection and localization by leveraging global contextual feature learning [4]. Unlike conventional Convolu-tional Neural Networks (CNNs), Vision Transformers capture long-range dependencies through self-attention mechanisms, enabling more effective identification of manipulated regions [4, 9]. The proposed framework performs image preprocessing, patch extraction, positional encod-ing, transformer-based feature learning, binary classification, and forgery localization. The model is evaluated using publicly available benchmark datasets, including CASIA V2, Co-MoFoD, and FaceForensics++ [20, 48], and its performance is assessed using Accuracy, Pre-cision, Recall, F1-score, Area Under Curve (AUC), Intersection over Union (IoU), and Pixel Accuracy [17, 49]. Experimental results demonstrate that the proposed Vision Transformer framework outperforms conventional CNN-based methods in terms of detection accuracy and localization precision [16, 19]. The proposed approach provides a robust and scalable solution for modern digital image forensics [15] and can be extended to hybrid transformer architectures and video forgery detection in future work.
Downloads
References
B. Bayar and M. C. Stamm, “A Deep Learning Approach to Universal Image Ma-nipulation Detection Using a New Convolutional Layer,” Proceedings of the ACM Workshop on Information Hiding and Multimedia Security, pp. 5–10, 2016. Available: https://scholar.google.com/scholar?q=Bayar+Stamm+image+manipulation+detection
D. Cozzolino, G. Poggi, and L. Verdoliva, “Recasting Residual-Based Local Descriptors as Convolutional Neural Networks,” IEEE Signal Processing Letters, vol. 24, no. 4, pp. 365–369, 2017. doi: 10.1109/LSP.2017.2651421
H. Dang, F. Liu, J. Stehouwer, X. Liu, and A. K. Jain, “On the Detec-tion of Digital Face Manipulation,” Proceedings of the IEEE/CVF Confer-ence on Computer Vision and Pattern Recognition (CVPR), 2020. Available: https://scholar.google.com/scholar?q=On+the+Detection+of+Digital+Face+Manipulation
A. Dosovitskiy et al., “An Image is Worth 16×16 Words: Transformers for Image Recog-nition at Scale,” International Conference on Learning Representations (ICLR), 2021.Available: https://scholar.google.com/scholar?q=An+Image+is+Worth+16x16+Words
I. Goodfellow, Y. Bengio, and A. Courville, Deep Learning. MIT Press, 2016. Available: https://scholar.google.com/scholar?q=Deep+Learning+Goodfellow
I. Goodfellow et al., “Generative Adversarial Nets,” Advances in Neural Information Processing Systems (NeurIPS), 2014. Available: https://scholar.google.com/scholar?q=Generative+Adversarial+Nets
Y. LeCun, Y. Bengio, and G. Hinton, “Deep Learning,” Nature, vol. 521, no. 7553, pp. 436–444, 2015. doi: 10.1038/nature14539
Y. Li and S. Lyu, “Exposing DeepFake Videos by Detect-ing Face Warping Artifacts,” CVPR Workshops, 2019. Available: https://scholar.google.com/scholar?q=Exposing+DeepFake+Videos
Z. Liu et al., “Swin Transformer: Hierarchical Vision Transformer Using Shifted Windows,” Proceedings of ICCV, 2021. Available: https://scholar.google.com/scholar?q=Swin+Transformer
T. M. Mitchell, Machine Learning. McGraw-Hill, 1997. Available: https://scholar.google.com/scholar?q=Mitchell+Machine+Learning
H. Nguyen, F. Fang, J. Yamagishi, and I. Echizen, “Multi-task Learning for Detect-ing and Segmenting Manipulated Facial Images and Videos,” IEEE International Con-ference on Biometrics, 2019. Available: https://scholar.google.com/scholar?q=Multi-task+Learning+for+Detecting+Manipulated+Facial+Images
A. Rossler et al., “FaceForensics++: Learning to Detect Manipulated Facial Im-ages,” IEEE Transactions on Pattern Analysis and Machine Intelligence, 2020. doi: 10.1109/TPAMI.2020.3001028
R. Salloum, Y. Ren, and C. C. J. Kuo, “Image Splicing Localization Using a Multi-task Fully Convolutional Network,” IEEE Transactions on Information Forensics and Security, 2018. Available: https://scholar.google.com/scholar?q=Image+Splicing+Localization
J. Schmidhuber, “Deep Learning in Neural Networks: An Overview,” Neural Networks, vol. 61, pp. 85–117, 2015. doi: 10.1016/j.neunet.2014.09.003
L. Verdoliva, “Media Forensics and DeepFakes: An Overview,” IEEE Journal of Selected Topics in Signal Processing, vol. 14, no. 5, pp. 910–932, 2020. doi: 10.1109/JSTSP.2020.3002101
C. Wang, X. Wu, and Z. Wang, “Image Forgery Detection Using Convolu-tional Neural Networks,” IEEE Access, vol. 7, pp. 85444–85455, 2019. Available: https://scholar.google.com/scholar?q=Image+Forgery+Detection+CNN
Y. Wu, W. Abd-Almageed, and P. Natarajan, “ManTra-Net: Manipulation Tracing Net-work for Detection and Localization of Image Forgeries,” Proceedings of CVPR, 2019.
Available: https://scholar.google.com/scholar?q=ManTra-Net
Y. Zhao et al., “ForgeryNet: A Versatile Benchmark for Comprehensive Forgery Analy-sis,” IEEE Transactions on Pattern Analysis and Machine Intelligence, 2022. Available: https://scholar.google.com/scholar?q=ForgeryNet
P. Zhou, X. Han, V. I. Morariu, and L. S. Davis, “Learning Rich Features for Im-age Manipulation Detection,” Proceedings of CVPR, pp. 1053–1061, 2018. Available:
https://scholar.google.com/scholar?q=Learning+Rich+Features+for+Image+Manipulation+Detection
P. Zhou et al., “FaceForensics++: Learning to Detect Manip-ulated Facial Images,” Proceedings of ICCV, 2019. Available: https://scholar.google.com/scholar?q=FaceForensics++
D. G. Lowe, “Distinctive Image Features from Scale-Invariant Keypoints,” Inter-national Journal of Computer Vision, vol. 60, no. 2, pp. 91–110, 2004. doi: 10.1023/B:VISI.0000029664.99615.94
H. Bay, A. Ess, T. Tuytelaars, and L. Van Gool, “Speeded-Up Robust Features (SURF),” Computer Vision and Image Understanding, vol. 110, no. 3, pp. 346–359, 2008. doi: 10.1016/j.cviu.2007.09.014
J. Fridrich, D. Soukal, and J. Lukas, “Detection of Copy-Move Forgery in Digital Images,” Proceedings of Digital Forensic Research Workshop, 2003.
A. C. Popescu and H. Farid, “Exposing Digital Forgeries by Detecting Duplicated Im-age Regions,” Department of Computer Science, Dartmouth College, Technical Report TR2004-515.
H. Farid, “Image Forgery Detection,” IEEE Signal Processing Magazine, vol. 26, no. 2,pp. 16–25, 2009. doi: 10.1109/MSP.2008.931079
J. Lukas, J. Fridrich, and M. Goljan, “Digital Camera Identification from Sensor Pattern Noise,” IEEE Transactions on Information Forensics and Security, vol. 1, no. 2, pp. 205–214, 2006. doi: 10.1109/TIFS.2006.873602
A. Swaminathan, M. Wu, and K. J. R. Liu, “Digital Image Forensics via Intrinsic Fin-gerprints,” IEEE Transactions on Information Forensics and Security, vol. 3, no. 1, pp. 101–117, 2008. doi: 10.1109/TIFS.2007.916285
B. Mahdian and S. Saic, “A Bibliography on Blind Methods for Identifying Image Forgery,” Signal Processing: Image Communication, vol. 25, pp. 389–399, 2010. doi: 10.1016/j.image.2010.04.001
H. Farid, “Digital Image Ballistics from JPEG Quantization,” Department of Computer Science, Dartmouth College, 2006.
A. C. Popescu and H. Farid, “Exposing Digital Forgeries in Color Filter Array Interpo-lated Images,” IEEE Transactions on Signal Processing, vol. 53, no. 10, pp. 3948–3959, 2005. doi: 10.1109/TSP.2005.855406
Y. LeCun, L. Bottou, Y. Bengio, and P. Haffner, “Gradient-Based Learning Applied to Document Recognition,” Proceedings of the IEEE, vol. 86, no. 11, pp. 2278–2324, 1998. doi: 10.1109/5.726791
A. Krizhevsky, I. Sutskever, and G. E. Hinton, “ImageNet Classification with Deep Con-volutional Neural Networks,” Advances in Neural Information Processing Systems, 2012. doi: 10.1145/3065386
K. He, X. Zhang, S. Ren, and J. Sun, “Deep Residual Learning for Image Recognition,” Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2016. doi: 10.1109/CVPR.2016.90
G. Huang, Z. Liu, L. Van Der Maaten, and K. Q. Weinberger, “Densely Connected Convo-lutional Networks,” Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2017. doi: 10.1109/CVPR.2017.243
F. Chollet, “Xception: Deep Learning with Depthwise Separable Convolutions,” Pro-ceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2017. doi: 10.1109/CVPR.2017.195
M. Tan and Q. Le, “EfficientNet: Rethinking Model Scaling for Convolutional Neural Networks,” Proceedings of the International Conference on Machine Learning, 2019.
Available: https://arxiv.org/abs/1905.11946
K. Simonyan and A. Zisserman, “Very Deep Convolutional Networks for Large-Scale Image Recognition,” International Conference on Learning Representations, 2015. Avail-able: https://arxiv.org/abs/1409.1556
C. Szegedy et al., “Going Deeper with Convolutions,” Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2015. doi: 10.1109/CVPR.2015.7298594
P. Zhou, X. Han, V. I. Morariu, and L. S. Davis, “Learning Rich Features for Image Manipulation Detection,” International Journal of Computer Vision, 2021.
A. Rossler et al., “FaceForensics++: Learning to Detect Manipulated Facial Images,” Proceedings of the IEEE International Conference on Computer Vision, 2019.
A. Vaswani et al., “Attention Is All You Need,” Advances in Neural Information Process-ing Systems (NeurIPS), 2017. doi: 10.5555/3295222.3295349
H. Touvron et al., “Training Data-Efficient Image Transformers and Distillation Through Attention,” Proceedings of the International Conference on Machine Learning (ICML), 2021.
N. Carion et al., “End-to-End Object Detection with Transformers,” European Conference on Computer Vision (ECCV), 2020.
C. Chen et al., “Vision Transformer for Image Recognition: A Survey,” IEEE Transac-tions on Pattern Analysis and Machine Intelligence, 2022.
S. Khan et al., “Transformers in Vision: A Survey,” ACM Computing Surveys, 2022. doi: 10.1145/3505244
J. Dong, W. Wang, and T. Tan, “CASIA Image Tampering Detection Evalua-tion Database,” Proceedings of the IEEE China Summit and International Confer-ence on Signal and Information Processing, pp. 422–426, 2013. doi: 10.1109/Chi-naSIP.2013.6625374
D. Tralic, J. Zupancic, S. Grgic, and M. Grgic, “CoMoFoD: New Database for Copy-Move Forgery Detection,” Proceedings of the 55th International Symposium ELMAR, pp. 49–54, 2013.
A. Rossler et al., “FaceForensics++: Learning to Detect Manipulated Facial Images,” IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 43, no. 10, pp. 3308–3320, 2020. doi: 10.1109/TPAMI.2020.3001028
Y. Zhao et al., “ForgeryNet: A Versatile Benchmark for Comprehensive Forgery Analy-sis,” IEEE Transactions on Pattern Analysis and Machine Intelligence, 2022.

This work is licensed under a Creative Commons Attribution 4.0 International License.
All articles published in our journal are licensed under CC-BY 4.0, which permits authors to retain copyright of their work. This license allows for unrestricted use, sharing, and reproduction of the articles, provided that proper credit is given to the original authors and the source.