Multimodal Fusion and Explainable Deep Learning for Synthetic Voice and Vishing Detection

Authors

Mahima BG

BMSCE, Dept of CSE (India)

Pallavi GB

BMSCE, Dept of CSE (India)

Article Information

DOI: 10.51244/IJRSI.2026.1307000015

Subject Category: Education

Volume/Issue: 13/7 | Page No: 218-229

Publication Timeline

Submitted: 2026-07-04

Accepted: 2026-07-09

Published: 2026-07-22

Abstract

Synthetic speech generation and voice-cloning technologies have achieved unprecedented levels of realism, enabling numerous applications in accessibility, virtual assistants, and media production. However, these advancements also introduce significant risks, including identity fraud, impersonation attacks, misinformation, and security breaches. This paper proposes a multimodal fusion framework for synthetic voice detection that combines handcrafted acoustic features with deep spectrogram representations to improve detection robustness and generalization. The proposed architecture employs a Convolutional Neural Network–Bidirectional Long ShortTerm Memory (CNN-BiLSTM) network to capture both spectral artifacts and temporal inconsistencies characteristic of AI-generated speech. To enhance transparency and interpretability, an explainability module incorporating attention visualization and feature attribution techniques is integrated into the detection pipeline. Furthermore, the framework is deployed through a real-time inference interface, demonstrating its practical applicability in cybersecurity, digital forensics, and media authentication scenarios. The findings highlight the effectiveness of combining deep learning, multimodal feature fusion, and explainable artificial intelligence to address the growing challenge of synthetic speech detection.

Keywords

Multimodal, Fusion, Explainable

Downloads

References

1. D. P. Kingma and J. Ba, “Adam: A method for stochastic optimization,” in Proc. International Conference on Learning Representations (ICLR), 2015. [Google Scholar] [Crossref]

2. A. Baevski, H. Zhou, A. Mohamed, and M. Auli, “wav2vec 2.0: A framework for self-supervised learning of speech representations,” in Proc. NeurIPS, 2020, pp. 12449–12460. [Google Scholar] [Crossref]

3. L. Breiman, “Random forests,” Machine Learning, vol. 45, no. 1, pp. 5–32, 2001. [Google Scholar] [Crossref]

4. K. Choi, G. Fazekas, M. Sandler, and K. Cho, “Convolutional recurrent neural networks for music classification,” in Proc. IEEE ICASSP, 2017, pp. 2392–2396. [Google Scholar] [Crossref]

5. C. Cortes and V. Vapnik, “Support-vector networks,” Machine Learning, vol. 20, no. 3, pp. 273–297, 1995. [Google Scholar] [Crossref]

6. H. Delgado, M. Todisco, M. Sahidullah, N. Evans, T. Kinnunen, K.-A. Lee, and J. Yamagishi, “ASVspoof 2021: Automatic speaker verification spoofing and countermeasures challenge evaluation plan,” 2021. [Google Scholar] [Crossref]

7. N. Evans, T. Kinnunen, and J. Yamagishi, “Spoofing and countermeasures for automatic speaker verification,” Speech Communication, vol. 66, pp. 130–153, 2015. [Google Scholar] [Crossref]

8. J. Frank and L. Schönherr, “WaveFake: A dataset to facilitate audio deepfake detection,” in Proc. Neural Information Processing Systems (NeurIPS) Workshop, 2021. [Google Scholar] [Crossref]

9. I. Goodfellow, Y. Bengio, and A. Courville, Deep Learning. Cambridge, MA, USA: MIT Press, 2016. [Google Scholar] [Crossref]

10. A. Graves and J. Schmidhuber, “Framewise phoneme classification with bidirectional LSTM and other neural network architectures,” Neural Networks, vol. 18, no. 5–6, pp. 602–610, 2005. [Google Scholar] [Crossref]

11. S. R. K. Khalid, S. Tariq, M. Kim, and S. S. Woo, “FakeAVCeleb: A novel audio-video multimodal deepfake dataset,” in Proc. NeurIPS Datasets and Benchmarks Track, 2021. [Google Scholar] [Crossref]

12. T. Kinnunen, M. Sahidullah, H. Delgado, M. Todisco, N. Evans, J. Yamagishi, and K.-A. Lee, “The ASVspoof 2017 challenge: Assessing the limits of replay spoofing attack detection,” in Proc. Interspeech, 2017, pp. 2–6. [Google Scholar] [Crossref]

13. A. Krizhevsky, I. Sutskever, and G. E. Hinton, “ImageNet classification with deep convolutional neural networks,” Communications of the ACM, vol. 60, no. 6, pp. 84–90, 2017. [Google Scholar] [Crossref]

14. M. Lundberg and S.-I. Lee, “A unified approach to interpreting model predictions,” in Proc. NeurIPS, 2017, pp. 4765–4774. [Google Scholar] [Crossref]

15. M. Mirsky and W. Lee, “The creation and detection of deepfakes: A survey,” ACM Computing Surveys, vol. 54, no. 1, pp. 1–41, 2021. [Google Scholar] [Crossref]

16. T. N. Sainath and C. Parada, “Convolutional neural networks for small-footprint keyword spotting,” in Proc. Interspeech, 2015, pp. 1478–1482. [Google Scholar] [Crossref]

17. K. N. Hsu, B. Bolte, Y.-H. H. Tsai, K. Lakhotia, R. Salakhutdinov, and A. Mohamed, “HuBERT: Self-supervised speech representation learning by masked prediction of hidden units,” IEEE/ACM Transactions on Audio, Speech, and Language Processing, vol. 29, pp. 3451–3460, 2021. [Google Scholar] [Crossref]

18. R. R. Selvaraju, M. Cogswell, A. Das, R. Vedantam, D. Parikh, and D. Batra, “Grad-CAM: Visual explanations from deep networks via gradient-based localization,” in Proc. IEEE ICCV, 2017, pp. 618–626. [Google Scholar] [Crossref]

19. S. Tolosana, R. Vera-Rodriguez, J. Fierrez, A. Morales, and J. Ortega-Garcia, “Deepfakes and beyond: A survey of face manipulation and fake detection,” Information Fusion, vol. 64, pp. 131–148, 2020. [Google Scholar] [Crossref]

20. M. Todisco, H. Delgado, and N. Evans, “A new feature for automatic speaker verification anti-spoofing: Constant Q cepstral coefficients,” in Proc. Odyssey Speaker and Language Recognition Workshop, 2016, pp. 283–290. [Google Scholar] [Crossref]

21. A. Vaswani et al., “Attention is All You Need,” in Proc. Advances in Neural Information Processing Systems (NeurIPS), 2017, pp. 5998–6008. [Google Scholar] [Crossref]

22. X. Wang, J. Yamagishi, M. Todisco, H. Delgado, N. Evans, T. Kinnunen, K.-A. Lee, V. Vestman, A. Nautsch, X. Qin, S. Sahidullah, J. Lorenzo-Trueba, and N. Dehak, “ASVspoof 2019: A large-scale public database of synthesized, converted and replayed speech,” Computer Speech & Language, vol. 64, Nov. 2020. [Google Scholar] [Crossref]

23. J. Yamagishi, X. Wang, M. Todisco, M. Sahidullah, H. Delgado, X. Qin, N. Evans, T. Kinnunen, K.-A. Lee, and V. Vestman, “ASVspoof 2021: Accelerating progress in spoofed and deepfake speech detection,” in Proc. ASVspoof Challenge Workshop, 2021. [Google Scholar] [Crossref]

24. T. Atmaja and M. Akagi, “Speech emotion recognition using deep learning methods: A review,” Electronics, vol. 10, no. 9, pp. 1–24, 2021. [Google Scholar] [Crossref]

25. D. Y. Mirsky and W. Lee, “The creation and detection of deepfakes: A survey,” ACM Computing Surveys, vol. 54, no. 1, pp. 1–41, 2021. [Google Scholar] [Crossref]

Metrics

Views & Downloads

Similar Articles