">

A Multimodel Deep Learning Framework for Detection of AI-Generated Images, Videos and Text



EOI: 10.11242/viva-tech.01.09.40

Download Full Text here



Citation

Karan Rout, Samarth More, Ved Parab, Prof. Minakshi Gaonkar,Yash Parab” A Multimodel Deep Learning Framework for Detection of AI-Generated Images, Videos and Text ", VIVA-IJRI Volume 1, Issue 9, Article 1, pp. 1-8, 2026. Published by Artificial Intelligence And Machine Learning Engineering Department, VIVA Institute of Technology, Virar, India.

Abstract

The rapid advancement of generative artificial intelligence has enabled the creation of highly realistic synthetic images, videos, and textual content, posing serious challenges to digital authenticity and information trust. While existing detection approaches largely focus on individual media modalities, real-world misinformation often involves a combination of visual and textual content. This paper presents a multimodal deep learning–based framework for detecting AI-generated synthetic media across images, videos, and text. The proposed approach integrates spatial feature analysis for images, temporal inconsistency detection for videos, and contextual language modelling for text to enable comprehensive content verification. Experiments are conducted using publicly available benchmark datasets, and the performance of the framework is evaluated using standard classification metrics. The results demonstrate that combining multimodal cues improves detection robustness compared to single-modality approaches. The findings highlight the effectiveness of multimodal analysis for addressing the growing challenges posed by AI-generated synthetic media in real-world digital environments.

Keywords

- Context-aware text analysis, Deep learning framework, Multimodal content verification, Synthetic media detection, Temporal inconsistency analysis, Visual artifact extraction.

References

  1. Rössler, Andreas, Davide Cozzolino, Luisa Verdoliva, Christian Riess, Justus Thies, and Matthias Nießner. “FaceForensics++: Learning to Detect Manipulated Facial Images.” In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV) (2019)
  2. Dolhansky, Brian, Russell Howes, Ben Pflaum, Nati Baram, and Cristian Canton Ferrer. “The DeepFake Detection Challenge (DFDC) Dataset.” arXiv preprint arXiv:2006.07397 (2020).
  3. Shu, Kai, Amy Sliva, Suhang Wang, Jiliang Tang, and Huan Liu. “Fake News Detection on Social Media: A Data Mining Perspective.” ACM SIGKDD Explorations Newsletter 19, no. 1 (2017): 22–36.
  4. Jawahar, Ganesh, Benoît Sagot, and Djamé Seddah. “What Does BERT Learn About the Structure of Language?” In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics (ACL) (2019).
  5. He, Kaiming, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. “Deep Residual Learning for Image Recognition.” In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR) (2016): 770–778.
  6. Chollet, François. “Xception: Deep Learning with Depthwise Separable Convolutions.” In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR) (2017): 1800–1807.
  7. Devlin, Jacob, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. “BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding.” In Proceedings of NAACL-HLT (2019): 4171–4186..
  8. Liu, Yinhan, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, et al. “RoBERTa: A Robustly Optimized BERT Pretraining Approach.” arXiv preprint arXiv:1907.11692 (2019).
  9. Agarwal, Shruti, Hany Farid, Yuming Gu, Mingming He, Koki Nagano, and Hao Li. “Protecting World Leaders Against Deep Fakes.” In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR) (2019): 38–45.
  10. Li, Yuezun, Ming-Ching Chang, and Siwei Lyu. “In Ictu Oculi: Exposing AI-Created Fake Videos by Detecting Eye Blinking.” In Proceedings of the IEEE International Workshop on Information Forensics and Security (WIFS) (2018): 1–7.
  11. Zhou, Peng, Xintong Han, Vlad I. Morariu, and Larry S. Davis. “Two-Stream Neural Networks for Tampered Face Detection.” In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition Workshops (CVPRW) (2017): 1831–1839.
  12. Afchar, Darius, Vincent Nozick, Junichi Yamagishi, and Isao Echizen. “MesoNet: A Compact Facial Video Forgery Detection Network.” In Proceedings of the IEEE International Workshop on Information Forensics and Security (WIFS) (2018): 1–7.
  13. Güera, David, and Edward J. Delp. “Deepfake Video Detection Using Recurrent Neural Networks.” In Proceedings of the IEEE International Conference on Advanced Video and Signal Based Surveillance (AVSS) (2018): 1–6.
  14. Nguyen, Huy H., Junichi Yamagishi, and Isao Echizen. “Use of a Capsule Network to Detect Fake Images and Videos.” In Proceedings of the IEEE International Conference on Image Processing (ICIP) (2019): 230–234.
  15. Wang, Sheng-Yu, Oliver Wang, Richard Zhang, Andrew Owens, and Alexei A. Efros. “CNN-Generated Images Are Surprisingly Easy to Spot… For Now.” In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) (2020): 8695–8704.
  16. Verdoliva, Luisa. “Media Forensics and Deepfakes: An Overview.” IEEE Journal of Selected Topics in Signal Processing 14, no. 5 (2020): 910–932.
  17. Zhou, Xinyi, Xiaohui Li, and Deepak Puthal. “Fake News Detection: A Survey.” Information Fusion 73 (2021): 1–21.
  18. Ruchansky, Natali, Sungyong Seo, and Yan Liu. “CSI: A Hybrid Deep Model for Fake News Detection.” In Proceedings of the ACM Conference on Information and Knowledge Management (CIKM) (2017): 797–806.
  19. Yang, Zichao, Diyi Yang, Chris Dyer, Xiaodong He, Alex Smola, and Eduard Hovy. “Hierarchical Attention Networks for Document Classification.” In Proceedings of NAACL-HLT (2016): 1480–1489.
  20. Sabir, Ekraam, Jia Cheng, Ayush Jaiswal, Wael AbdAlmageed, Iacopo Masi, and Prem Natarajan. “Recurrent Convolutional Strategies for Face Manipulation Detection in Videos.” In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops (CVPRW) (2019).
  21. Mittal, Trisha, Utkarsh Bhattacharya, Rohan Chandra, Aniket Bera, and Dinesh Manocha. “Emotions Don’t Lie: An Audio-Visual Deepfake Detection Method Using Affective Cues.” In Proceedings of the ACM International Conference on Multimedia (2020): 2823–2832.
  22. Ganguly, Debasis, and Gareth J. F. Jones. “Dynamic Word Embeddings for Evolving Semantic Discovery.” ACM Transactions on Information Systems 36, no. 4 (2018): 1–34.
  23. Li, Yuezun, Xin Yang, Peng Sun, Hao Qi, and Siwei Lyu. “Celeb-DF: A Large-Scale Challenging Dataset for Deepfake Forensics.” In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) (2020): 3207–3216.
  24. Zhao, Hang, Weiyang Zhou, Chao Chen, and Hao Li. “Multi-Attentional Deepfake Detection.” In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) (2021): 2185–2194.
  25. Haliassos, Alexandros, Mihalis Miron, Stavros Petridis, and Maja Pantic. “Lip Reading Deepfake Detection.”
  26. Chen, L., Zhao, Y., & Kumar, S. (2025). Multimodal misinformation detection using cross-modal transformers. IEEE Transactions on Multimedia, 27, 1123–1136.
  27. Park, J., Singh, R., & Lee, H. (2025). Vision-language foundation models for synthetic media verification. Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR).
  28. Ahmed, T., Wang, Z., & Li, Q. (2026). Robust multimodal fake news detection under adversarial settings. Information Fusion, 95, 214–228.