Image Authentication: Differentiating Camera-Captured and AI-Generated Images with Provenance Verification using Vision Transformer (ViT)
Authors
Department of Computer Science, GITAM School of Science, GITAM (Deemed to be University), Visakhapatnam (India)
Article Information
DOI: 10.51584/IJRIAS.2026.11060074
Subject Category: Artificial Intelligence
Volume/Issue: 11/6 | Page No: 835-846
Publication Timeline
Submitted: 2026-06-04
Accepted: 2026-06-10
Published: 2026-06-23
Abstract
The rapid growth of generative artificial intelligence significantly affects the development of digital image creation. Recently, advances in Generative Adversarial Networks (GAN) and Diffusion-based Generative Networks have made it possible to create synthetic images that are hard to tell apart from real ones. Therefore, verifying images produced by artificial intelligence is essential. This poses a major challenge for many fields, including digital forensics, security, and image verification. There has been a rise in the misuse of AI-generated images to spread false news, impersonate people, and manipulate images. Traditional methods for verifying image authenticity, such as human observation and image metadata analysis, are now unreliable. AI-generated images can be easily altered, and human observation alone cannot confirm an image's authenticity. As a result, there is a pressing need to develop an effective image authentication system to tell apart real images from those created by AI.
This paper proposes an automated image authentication system. It utilizes a Vision Transformer model for classifying images as either camera-captured or AI-generated. The system employs a pre-trained model for feature extraction, which is then fine-tuned for classification. Unlike conventional convolutional neural networks, the Vision Transformer treats an image as a sequence of patches and uses self-attention to capture global dependencies. This method helps to identify subtle differences in AI-generated images. Additionally, the proposed system incorporates a confidence level represented by Softmax probabilities, which helps understand the reliability of the system's results. An explainability feature is also included, using Explainable Artificial Intelligence techniques to highlight areas in the image that influence the results. This system provides a strong solution for the current challenges in image authentication. It can be implemented as a web-based application using the Flask framework. Experimental results demonstrate that the system achieves high accuracy in classifying AI-generated images.
Keywords
Image Authentication, Vision Transformer, AI generated images
Downloads
References
1. D. Karageorgiou et al., “Any-resolution AI-generated image detection by spectral learning,” arXiv preprint, 2025. [Google Scholar] [Crossref]
2. Y. Zhang et al., “Unmasking AI-created visual content: A review of generated images and deepfake detection technologies,” Journal of King Saud University - Computer and Information Sciences, 2025. [Google Scholar] [Crossref]
3. S. Tang et al., “Towards extensible detection of AI-generated images via content-agnostic adapter-based category-aware incremental learning,” IEEE Transactions on Information Forensics and Security, 2025. [Google Scholar] [Crossref]
4. J. J. Bird and A. Lotfi, “CIFAKE: Image classification and explainable identification of AI-generated synthetic images,” IEEE Access, vol. 12, pp. 156379–156393, 2024. [Google Scholar] [Crossref]
5. N. Carlini et al., “Extracting training data from diffusion models,” in Proc. USENIX Security Symposium, 2023. [Google Scholar] [Crossref]
6. R. Rombach et al., “High-resolution image synthesis with latent diffusion models,” in Proc. IEEE/CVF Conf. on Computer Vision and Pattern Recognition (CVPR), 2022, pp. 10684–10695. [Google Scholar] [Crossref]
7. A. Ramesh et al., “Hierarchical text-conditional image generation with CLIP latents,” OpenAI, 2022. [Google Scholar] [Crossref]
8. Z. Liu et al., “Swin transformer: Hierarchical vision transformer using shifted windows,” in Proc. IEEE/CVF Int. Conf. on Computer Vision (ICCV), 2021, pp. 10012–10022. [Google Scholar] [Crossref]
9. X. Zhang et al., “Detecting AI-synthesized images using frequency domain analysis,” IEEE Access, vol. 9, pp. 124256–124268, 2021. [Google Scholar] [Crossref]
10. B. Tondi et al., “Detection of GAN-generated fake images over social networks,” in Proc. IEEE ICASSP, 2021, pp. 3010–3014. [Google Scholar] [Crossref]
11. J. Frank et al., “Leveraging frequency analysis for deep fake image recognition,” in Proc. ICML Workshops, 2020. [Google Scholar] [Crossref]
12. S. Wang et al., “CNN-generated images are surprisingly easy to spot… for now,” in Proc. IEEE/CVF CVPR, 2020, pp. 8695–8704. [Google Scholar] [Crossref]
13. H. Dang et al., “On the detection of digital face manipulation,” in Proc. IEEE/CVF CVPR Workshops, 2020. [Google Scholar] [Crossref]
14. S. Verdoliva, “Media forensics and deepfakes: An overview,” IEEE Journal of Selected Topics in Signal Processing, vol. 14, no. 5, pp. 910–932, 2020. [Google Scholar] [Crossref]
15. P. Korshunov and S. Marcel, “Deepfake detection: A survey,” IEEE Signal Processing Magazine, vol. 36, no. 1, pp. 85–95, 2019. [Google Scholar] [Crossref]
16. Y. Nirkin et al., “Inverting face embeddings with convolutional neural networks,” in Proc. IEEE/CVF CVPR, 2019, pp. 1070–1078. [Google Scholar] [Crossref]
17. M. Tan and Q. Le, “EfficientNet: Rethinking model scaling for convolutional neural networks,” in Proc. ICML, 2019, pp. 6105–6114. [Google Scholar] [Crossref]
18. H. Nguyen et al., “Capsule-forensics: Using capsule networks to detect forged images and videos,” in Proc. IEEE ICIP, 2019, pp. 2307–2311. [Google Scholar] [Crossref]
19. M. Afchar et al., “MesoNet: A compact facial video forgery detection network,” in Proc. IEEE WIFS, 2018, pp. 1–7. [Google Scholar] [Crossref]
20. F. Chollet, “Xception: Deep learning with depthwise separable convolutions,” in Proc. IEEE CVPR, 2017, pp. 1251–1258. [Google Scholar] [Crossref]
Metrics
Views & Downloads
Similar Articles
- The Role of Artificial Intelligence in Revolutionizing Library Services in Nairobi: Ethical Implications and Future Trends in User Interaction
- ESPYREAL: A Mobile Based Multi-Currency Identifier for Visually Impaired Individuals Using Convolutional Neural Network
- Comparative Analysis of AI-Driven IoT-Based Smart Agriculture Platforms with Blockchain-Enabled Marketplaces
- AI-Based Dish Recommender System for Reducing Fruit Waste through Spoilage Detection and Ripeness Assessment
- SEA-TALK: An AI-Powered Voice Translator and Southeast Asian Dialects Recognition