Multimodal Depression Detection from Speech and Text Using a Fusion Neural Network
Authors
Department of Computer Science, Taraba State University, Jalingo, Nigeria (Nigeria)
Department of Computer Science, Taraba State University, Jalingo, Nigeria (Nigeria)
Department of Computer Science, Taraba State University, Jalingo, Nigeria (Nigeria)
Department of Computer Science, Taraba State University, Jalingo, Nigeria (Nigeria)
Article Information
DOI: 10.51244/IJRSI.2026.1308000043
Subject Category: Education
Volume/Issue: 13/8 | Page No: 515-527
Publication Timeline
Submitted: 2026-08-03
Accepted: 2026-08-08
Published: 2026-08-31
Abstract
Depression is among the leading causes of disability worldwide, yet its detection continues to rely heavily on subjective clinical interviews and self-report instruments that are difficult to scale, particularly in low-resource regions. This paper presents the design, implementation, and empirical validation of a multimodal fusion neural network for automatic depression detection from speech and text, developed with specific attention to closing the near-total absence of African research contributions in this rapidly growing area of computing. The proposed system extracts spectral-prosodic acoustic features from voice recordings using a convolutional encoder, and contextual semantic features from transcribed text using a transformer-based language encoder, before an attention gate learns to weight the relative contribution of each modality per instance ahead of a fully connected classifier. The architecture was implemented and validated end-to-end as a feasibility study on two accessible, weak-label proxy corpora: the RAVDESS acted-emotion speech dataset, with sad and calm recordings relabeled as a depression-like class, and the dair-ai/emotion text corpus, with sadness and fear posts relabeled likewise. On held-out test data, the fusion model reached 90.28% accuracy and 0.9609 AUC-ROC, and, most notably, recovered the positive-class recall that the audio-only model lost almost entirely (21.05% versus 84.21%), demonstrating that the attention-gated fusion mechanism functions as designed. This paper reports the background, problem definition, related work, the complete mathematical formulation of the implemented pipeline, the empirical results of this proxy validation, and a discussion of what they do and do not establish, closing with concrete recommendations for advancing the work toward a clinically meaningful, Africa-relevant screening tool.
Keywords
Depression detection, multimodal fusion, deep learning, speech processing, natural language processing, attention mechanism, proxy-label validation, affective computing, mental health informatics.
Downloads
References
1. M. Fang, S. Peng, Y. Liang, C. C. Hung, and S. Liu, “A multimodal fusion model with multi-level attention mechanism for depression detection,” Biomedical Signal Processing and Control, vol. 82, p. 104561, 2023. [Google Scholar] [Crossref]
2. M. Nykoniuk, O. Basystiuk, N. Shakhovska, and N. Melnykova, “Multimodal data fusion for depression detection approach,” Computation, vol. 13, no. 1, p. 9, 2025. [Google Scholar] [Crossref]
3. Z. Xu, Y. Gao, F. Wang, L. Zhang, L. Zhang, J. Wang, and J. Shu, “Depression detection methods based on multimodal fusion of voice and text,” Scientific Reports, vol. 15, no. 1, p. 21907, 2025. [Google Scholar] [Crossref]
4. Y. Xia, L. Liu, T. Dong, J. Chen, Y. Cheng, and L. Tang, “A depression detection model based on multimodal graph neural network,” Multimedia Tools and Applications, vol. 83, no. 23, pp. 63379–63395, 2024. [Google Scholar] [Crossref]
5. Z. Zhang, S. Zhang, D. Ni, Z. Wei, K. Yang, S. Jin, et al., “Multimodal sensing for depression risk detection: Integrating audio, video, and text data,” Sensors, vol. 24, no. 12, p. 3714, 2024. [Google Scholar] [Crossref]
6. M. Rohanian, J. Hough, and M. Purver, “Detecting depression with word-level multimodal fusion,” in Proc. Interspeech 2019, 2019, pp. 1443–1447. [Google Scholar] [Crossref]
7. Y. Li, X. Yang, M. Zhao, J. Wang, Y. Yao, W. Qian, and S. Qi, “Predicting depression by using a novel deep learning model and video-audio-text multimodal data,” Frontiers in Psychiatry, vol. 16, p. 1602650, 2025. [Google Scholar] [Crossref]
8. J. Ye, Y. Yu, Q. Wang, W. Li, H. Liang, Y. Zheng, and G. Fu, “Multi-modal depression detection based on emotional audio and evaluation text,” Journal of Affective Disorders, vol. 295, pp. 904–913, 2021. [Google Scholar] [Crossref]
9. F. Mohammad and K. M. Al Mansoor, “MDD: A unified multimodal deep learning approach for depression diagnosis based on text and audio speech,” Computers, Materials & Continua, vol. 81, no. 3, pp. 4125–4147, 2024. [Google Scholar] [Crossref]
10. H. Yoo and H. Oh, “Depression detection model using multimodal deep learning,” 2023. [Google Scholar] [Crossref]
11. Y. H. Hu, R. Y. Wu, M. Y. Su, I. L. Lin, and C. C. Shen, “Multimodal multitask learning for predicting depression severity and suicide risk using pretrained audio and text embeddings: Methodology development and application,” JMIR Medical Informatics, vol. 13, p. e66907, 2025. [Google Scholar] [Crossref]
12. Y. Jin, X. Chen, J. Liu, Z. Chen, J. Zhou, Y. Li, et al., “Harnessing multimodal emotion features in depression detection across gender: Integrating large language model, acoustic fusion and facial expression recognition,” Journal of Affective Disorders, p. 121207, 2026. [Google Scholar] [Crossref]
13. L. Zhang, Y. Fan, J. Jiang, Y. Li, and W. Zhang, “Adolescent depression detection model based on multimodal data of interview audio and text,” International Journal of Neural Systems, vol. 32, no. 11, p. 2250045, 2022. [Google Scholar] [Crossref]
14. S. Mamidisetti and M. Reddy, “Multimodal depression detection using audio, visual and textual cues: A survey,” NeuroQuantology, vol. 20, no. 4, pp. 325–338, 2022. [Google Scholar] [Crossref]
15. A. Sharma, A. Saxena, A. Kumar, and D. Singh, “Depression detection using multimodal analysis with chatbot support,” in Proc. 2nd Int. Conf. Disruptive Technologies (ICDT), 2024, pp. 328–334. [Google Scholar] [Crossref]
16. H. Zhang, H. Wang, S. Han, W. Li, and L. Zhuang, “Detecting depression tendency with multimodal features,” Computer Methods and Programs in Biomedicine, vol. 240, p. 107702, 2023. [Google Scholar] [Crossref]
17. X. Zhang, B. Li, and G. Qi, “A novel multimodal depression diagnosis approach utilizing a new hybrid fusion method,” Biomedical Signal Processing and Control, vol. 96, p. 106552, 2024. [Google Scholar] [Crossref]
18. J. Chen, S. Liu, M. Xu, and P. Wang, “Enhancing depression detection: A multimodal approach with text extension and content fusion,” Expert Systems, vol. 41, no. 10, p. e13616, 2024. [Google Scholar] [Crossref]
19. Y. Zhang, Y. Wang, X. Wang, B. Zou, and H. Xie, “Text-based decision fusion model for detecting depression,” in Proc. 2nd Symp. Signal Processing Systems, 2020, pp. 101–106. [Google Scholar] [Crossref]
20. R. P. Thati, A. S. Dhadwal, P. Kumar, and P. Sainaba, “Multimodal depression detection: Using fusion strategies with smart phone usage and audio-visual behavior,” International Journal on Artificial Intelligence Tools, vol. 32, no. 02, p. 2340008, 2023. [Google Scholar] [Crossref]
21. L. Wang, Y. Zhang, B. Zhou, S. Cao, K. Hu, and Y. Tan, “Automatic depression prediction via cross-modal attention-based multi-modal fusion in social networks,” Computers and Electrical Engineering, vol. 118, p. 109413, 2024. [Google Scholar] [Crossref]
22. M. Rodrigues Makiuchi, T. Warnita, K. Uto, and K. Shinoda, “Multimodal fusion of BERT-CNN and gated CNN representations for depression detection,” in Proc. 9th Int. Audio/Visual Emotion Challenge and Workshop, 2019, pp. 55–63. [Google Scholar] [Crossref]
23. Y. Wang, Z. Wang, C. Li, Y. Zhang, and H. Wang, “A multimodal feature fusion-based method for individual depression detection on Sina Weibo,” in Proc. IEEE 39th Int. Performance Computing and Communications Conf. (IPCCC), 2020, pp. 1–8. [Google Scholar] [Crossref]
24. H. Naderi, B. H. Soleimani, and S. Matwin, “Multimodal deep learning for mental disorders prediction from audio speech samples,” arXiv preprint arXiv:1909.01067, 2019. [Google Scholar] [Crossref]
25. X. Zhang, X. Gong, W. Li, G. Liu, and Y. Li, “Depression detection using BiLSTM multi-head attention fusion network,” Expert Systems with Applications, p. 130100, 2025. [Google Scholar] [Crossref]
26. L. Yang, D. Jiang, X. Xia, E. Pei, M. C. Oveneke, and H. Sahli, “Multimodal measurement of depression using deep learning models,” in Proc. 7th Annual Workshop on Audio/Visual Emotion Challenge, 2017, pp. 53–59. [Google Scholar] [Crossref]
27. N. Wang, R. Chiong, R. Kamil, W. Zhang, S. A. R. Al-Haddad, and N. Ibrahim, “Depression detection using speech audio and text: A comprehensive review focusing on deep learning methods,” 2024. [Google Scholar] [Crossref]
28. A. Y. Kim, E. H. Jang, S. H. Lee, K. Y. Choi, J. G. Park, and H. C. Shin, “Automatic depression detection using smartphone-based text-dependent speech signals: Deep convolutional neural network approach,” Journal of Medical Internet Research, vol. 25, p. e34474, 2023. [Google Scholar] [Crossref]
29. W. Zhang, K. Mao, and J. Chen, “A multimodal approach for detection and assessment of depression using text, audio and video,” Phenomics, vol. 4, no. 3, pp. 234–249, 2024. [Google Scholar] [Crossref]
30. W. Xie, C. Wang, Z. Lin, X. Luo, W. Chen, M. Xu, et al., “Multimodal fusion diagnosis of depression and anxiety based on CNN-LSTM model,” Computerized Medical Imaging and Graphics, vol. 102, p. 102128, 2022. [Google Scholar] [Crossref]
31. H. Ding, Z. Du, Z. Wang, J. Xue, Z. Wei, K. Yang, et al., “IntervoxNet: A novel dual-modal audio-text fusion network for automatic and efficient depression detection from interviews,” Frontiers in Physics, vol. 12, p. 1430035, 2024. [Google Scholar] [Crossref]
32. F. F. D. Almeida, K. R. T. Aires, A. C. B. Soares, L. D. S. B. Neto, and R. D. M. S. Veras, “Multimodal fusion for depression detection assisted by stacking deep neural networks,” 2024. [Google Scholar] [Crossref]
33. M. He, E. M. Bakker, and M. S. Lew, “DPD (DePression Detection) Net: A deep neural network for multimodal depression detection,” Health Information Science and Systems, vol. 12, no. 1, p. 53, 2024. [Google Scholar] [Crossref]
34. Y. Shi, T. Gan, J. Li, and S. Li, “LSHMMformer: An intelligent detection model of depression based on multi-modal fusion,” Journal of Neuroscience Methods, p. 110755, 2026. [Google Scholar] [Crossref]
35. Z. Cheng, X. Huang, and Y. Ding, “An intelligent depression detection model based on multimodal fusion technology,” Journal of Mechanics in Medicine and Biology, vol. 24, no. 08, p. 2440046, 2024. [Google Scholar] [Crossref]
36. L. Yang, D. Jiang, and H. Sahli, “Integrating deep and shallow models for multi-modal depression analysis—Hybrid architectures,” IEEE Transactions on Affective Computing, vol. 12, no. 1, pp. 239–253, 2018. [Google Scholar] [Crossref]
37. Z. Zhang, W. Lin, M. Liu, and M. Mahmoud, “Multimodal deep learning framework for mental disorder recognition,” in Proc. 15th IEEE Int. Conf. Automatic Face and Gesture Recognition (FG), 2020, pp. 344–350. [Google Scholar] [Crossref]
38. L. Ilias and D. Askounis, “A cross-attention layer coupled with multimodal fusion methods for recognizing depression from spontaneous speech,” in Proc. Interspeech 2024, 2024, pp. 912–916. [Google Scholar] [Crossref]
39. S. Hemalatha, K. Jothimani, K. Swathi, S. Shibinta, W. Jason Selvakumar, and D. Sathish, “Multimodal approach for depression detection: Integrating speech and eye ball movement data,” in Proc. 1st Int. Conf. Advances in Computing, Communication and Networking (ICAC2N), 2024, pp. 1011–1019. [Google Scholar] [Crossref]
40. A. Patel, A. Kumar, M. Sharma, N. Pandey, and M. H. Gautam, “Multimodal AI-based mental health detection system: Integrating text, speech, and facial analysis using deep learning,” 2026. [Google Scholar] [Crossref]
41. J. Gratch, R. Artstein, G. Lucas, G. Stratou, S. Scherer, A. Nazarian, R. Wood, J. Boberg, D. DeVault, S. Marsella, D. Traum, S. Rizzo, and L.-P. Morency, “The distress analysis interview corpus of human and computer interviews,” in Proc. 9th Int. Conf. Language Resources and Evaluation (LREC), Reykjavik, Iceland, 2014, pp. 3123–3128. [Google Scholar] [Crossref]
42. G. Coppersmith, M. Dredze, C. Harman, K. Hollingshead, and M. Mitchell, “CLPsych 2015 shared task: Depression and PTSD on Twitter,” in Proc. 2nd Workshop on Computational Linguistics and Clinical Psychology: From Linguistic Signal to Clinical Reality, Denver, CO, USA, 2015, pp. 31–39. [Google Scholar] [Crossref]
43. F. Eyben, M. Wöllmer, and B. Schuller, “openSMILE: The Munich versatile and fast open-source audio feature extractor,” in Proc. 18th ACM Int. Conf. Multimedia, Florence, Italy, 2010, pp. 1459–1462. [Google Scholar] [Crossref]
44. M. Valstar, J. Gratch, B. Schuller, F. Ringeval, D. Lalanne, M. Torres Torres, S. Scherer, G. Stratou, R. Cowie, and M. Pantic, “AVEC 2016: Depression, mood, and emotion recognition workshop and challenge,” in Proc. 6th Int. Workshop on Audio/Visual Emotion Challenge, 2016, pp. 3–10. [Google Scholar] [Crossref]
45. M. De Choudhury, M. Gamon, S. Counts, and E. Horvitz, “Predicting depression via social media,” in Proc. 7th Int. AAAI Conf. Weblogs and Social Media (ICWSM), 2013, pp. 128–137. [Google Scholar] [Crossref]
Metrics
Views & Downloads
Similar Articles
- Assessment of the Role of Artificial Intelligence in Repositioning TVET for Economic Development in Nigeria
- Teachers’ Use of Assure Model Instructional Design on Learners’ Problem Solving Efficacy in Secondary Schools in Bungoma County, Kenya
- “E-Booksan Ang Kaalaman”: Development, Validation, and Utilization of Electronic Book in Academic Performance of Grade 9 Students in Social Studies
- Analyzing EFL University Students’ Academic Speaking Skills Through Self-Recorded Video Presentation
- Major Findings of The Study on Total Quality Management in Teachers’ Education Institutions (TEIs) In Assam – An Evaluative Study