Multimodal Deepfake Detection for Digital Forensics: A Robust Audio-Visual Inconsistency Approach for Evidence Integrity

Authors

  • Truong Xuan Hung Academy of Cryptography Techniques
  • Luong The Dung
  • Tran Anh Tu

DOI:

https://doi.org/10.54654/isj.v2i28.1248

Keywords:

Multimedia Forensics, deepfake detection, audioVisual synchronization, contrastive learning, eKYC security

Tóm tắt

The rapid proliferation of Generative AI (GenAI) has democratized the creation of hyper-realistic multimedia forgeries, posing severe threats to electronic Know Your Customer (eKYC) systems and digital forensic investigations. While visual synthesis has reached near-perfection, maintaining precise synchronization between lip movements (visemes) and speech signals (phonemes) remains a formidable challenge. To address this, we propose a novel Multimodal Deep Learning framework designed to detect high-fidelity Deepfakes by exploiting audio visual temporal inconsistencies. Beyond traditional feature fusion, our architecture integrates a Contrastive Synchronization Loss with a Transformer based Cross-Modal Attention mechanism. This hybrid objective explicitly enforces intra-class compactness for authentic pairs while amplifying the distance for asynchronous forgeries. Extensive experiments on FaceForensics++, DFDC, and a custom Vietnamese dataset (Vn-eKYC-Aug) demonstrate that our model achieves state-of-the-art performance, maintaining high robustness against video compression and environmental noise, though operational efficacy remains sensitive to extreme low-light conditions and diverse regional dialects. This research provides a resilient forensic layer for digital identity verification, ensuring evidence integrity in the GenAI era.

Downloads

Download data is not yet available.

References

Yu, Z., Qin, Y., Li, X., Zhao, C., Lei, Z., Zhao, G."Deep Learning for Face Anti-Spoofing: A Survey", in IEEE Transactions on Pattern Analysis and Machine Intelligence (TPAMI) 45(5), pp. 5609-5631, 2023.

Shao, R., Wu, T., Liu, Z."Detecting and Recovering Sequential DeepFake Manipulation", in Proceedings of the European Conference on Computer Vision (ECCV), 2024. DOI: 10.48550/arXiv.2207.02204.

Rombach, R., Blattmann, A., Lorenz, D., Esser, P., Ommer, B. "High-Resolution Image Synthesis with Latent Diffusion Models", in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). pp. 10684-10695, 2022.

Coccomini, D.A., Esuli, A., Falchi, F., Gennaro, C., Messina, N."Detecting images generated by deep diffusion models using their local intrinsic dimensionality". IEEE Access 11, pp. 106199–106211, 2023.

Rossler, ¨ A., Cozzolino, D., Verdoliva, L., Riess, C.,Thies, J., Nießner, M.: FaceForensics++. "Learning to Detect Manipulated Facial Images", in Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), pp. 1-11, 2019.

Todisco, M., Wang, X., Vestman, V., Sahidullah, M., Delgado, H., Nautsch, A., Yamagishi, J., Evans, N., Kinnunen, T., Lee, K.A. "ASVspoof 2019: Future Horizons in Spoofed and Fake Audio Detection", in Proceedings of the Annual Conference of the International Speech Communication Association (Interspeech), pp. 1008–1012 (2019).

Wang, Z., Bao, J., Zhou, W., Wang, W., Hu, H., Chen, H., Li, H. "AltFreezing for More General Video Face Forgery Detection", in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pp. 4129-4139, 2023.

Cocchi, F., Baraldi, L., Poppi, S., Cornia, M., Baraldi, L., Cucchiara, R. "Unveiling the Impact of Image Transformations on Deepfake Detection: An Experimental Analysis", in Proceedings of the International Conference on Image Analysis and Processing (ICIAP), pp. 345-356, 2023.

Cheng, Y., Guo, M., Zhao, X., Liang, J., Wang, D."Voice-Face Homogeneity Tells Deepfake Videos", in IEEE International Conference on Multimedia and Expo (ICME), pp. 67-72, 2023.

Nguyen, T.T., Nguyen, Q.V.H., Nguyen, D.T., Nguyen, D.T., Huynh-The, T., Nahavandi, S., Nguyen, T.T., Pham, Q.-V., Nguyen, C.M."Deep learning for deepfakes creation and detection: A survey",in Computer Vision and Image Understanding 223, 103525, 2022. DOI: 10.1016/j.cviu.2022.103525.

Nguyen, T.T.T., Nguyen, H.K. "VLSP 2021 - TTS Challenge: Vietnamese Spontaneous Speech Synthesis", VNU Journal of Science: Computer Science and Communication Engineering vol 38 no 1, pp. 25-33, 2022

Pham, V.H., Do, T.T.H. "Enhancing Web Application Security: A Deep Learning and NLP-based Approach for Accurate Attack Detection", Journal of Science and Technology on Information Security, vol. 3, no. 20, pp. 77-87. DOI: 10.54654/isj.v3i20.1008.

Tong, A.T., Nguyen, N.C., Nguyen, V.A., Hoang, V.L. "Proposing the application of a deep learning model to detect the malicious IP address of botnet in the computer network", Journal of Science and Technology on Information Security vol.3, no.17, pp. 43-52. DOI: 10.54654/isj.v3i17.894.

Masood, M., Nawaz, M., Malik, K., Javed, A., Irtaza, A., Malik, H."Deepfakes generation and detection: state-of-the-art, open challenges, countermeasures, and way forward", in Applied Intelligence. vol. 53, pp. 1-53, 2022.

Guera, ¨ D., Delp, E.J. "Deepfake Video Detection Using Recurrent Neural Networks" in Proceedings of the 15th IEEE International Conference on Advanced Video and Signal Based Surveillance (AVSS). pp. 1–6, 2018.

Haliassos, A., Vougioukas, K., Petridis, S., Pantic, M. "Lips Don’t Lie: A Generalisable and Robust Approach to Face Forgery Detection", in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pp. 5037-5047, 2021).

Afchar, D., Nozick, V., Yamagishi, J., Echizen, I."MesoNet: a Compact Facial Video Forgery Detection Network", in Proceedings of the IEEE International Workshop on Information Forensics and Security (WIFS). pp. 1-7, 2018.

Liu, Z., Lin, Y., Cao, Y., Hu, H., Wei, Y., Zhang, Z., Lin, S., Guo, B. "Swin Transformer: Hierarchical Vision Transformer using Shifted Windows", in Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), 2021

Durall, R., Keuper, M., Keuper, J. "Watch Your UpConvolution: CNN Based Generative Color Naming Exposes Fake Images", in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). pp. 4941-4950, 2020.

Cozzolino, D., Poggi, G., Corvi, R., Nießner, M., Verdoliva, L. "Raising the Bar of AI-generated Image Detection with CLIP", in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops (CVPRW), pp. 4356-4366, 2024.

H. Zhao, T. Wei, W. Zhou, W. Zhang, D. Chen and N. Yu, "Multi-attentional Deepfake Detection", IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pp. 2185-2194, Nashville, TN, USA, 2021, DOI: 10.1109/CVPR46437.2021.00222.

Haliassos, A., Mira, R., Petridis, S., Pantic, M."Leveraging Real Talking Faces via Self-Supervision for Robust Forgery Detection" in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pp. 14950-14962, 2022.

Ma, P., Haliassos, A., Fernandez-Lopez, A., Lu, H., Petridis, S., Pantic, M. "AV-Deepfake1M: A LargeScale Audio-Visual Deepfake Dataset and A Unified Baseline", in Thirty-seventh Conference on Neural Information Processing Systems (NeurIPS) Datasets and Benchmarks Track, 2023.

Chung, J.S., Senior, A., Vinyals, O., Zisserman, A. "Lip Reading Sentences in the Wild", in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pp. 3444-3453, 2017.

Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A.N., Kaiser, Ł., Polosukhin, I. "Attention is all you need", in Advances in Neural Information Processing Systems (NeurIPS), pp. 6000-6010, 2017.

Chung, J.S., Zisserman, A. "Out of time: automated lip sync in the wild", in Asian Conference on Computer Vision (ACCV), 2016.

Dolhansky, B., Bitton, J., Pflaum, B., Lu, J., Howes, R., Wang, M., Ferrer, C. "The DeepFake Detection Challenge Dataset", in Computer Vision and Pattern Recognition (cs.CV), 2020. DOI:10.48550/arXiv.2006.07397.

Zhou, Y., Lim, S.N. "Joint Audio-Visual Deepfake Detection", in Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), pp. 14800-14809, 2021.

Downloads

Abstract views: 25 / PDF downloads: 7

Published

2026-08-16

How to Cite

Hung, T. X., Dung, L. T., & Tu, T. A. (2026). Multimodal Deepfake Detection for Digital Forensics: A Robust Audio-Visual Inconsistency Approach for Evidence Integrity. Journal of Science and Technology on Information Security, 2(28), 5-26. https://doi.org/10.54654/isj.v2i28.1248

Issue

Section

Papers

Most read articles by the same author(s)