A Real-Time Multimodal Approach to Mental Health Monitoring and Analysis

Authors

Mrs. T. Gayathri Devi

Assistant Professor, Department of Information Technology, MNM Jain Engineering College, Chennai, India (India)

Article Information

DOI: 10.51584/IJRIAS.2026.11060013

Subject Category: Artificial Intelligence

Volume/Issue: 11/6 | Page No: 111-116

Publication Timeline

Submitted: 2026-05-31

Accepted: 2026-06-05

Published: 2026-06-17

Abstract

The Multimodal Mental Health Analysis System collects inputs including PHQ-9questionnaire responses, written personal narratives, spoken video recordings, and optional location data to assess a user's mental wellness through a combined analysis pipeline. Text responses are processed using scoring methods and natural language processing techniques to identify patterns related to emotional distress, depressive symptoms, anxiety, burnout, and overall mental health indicators. Audio extracted from the spoken video response is analysed for transcript content, pitch, energy, and vocal variation to capture tone-related cues, while video frames are examined using computer vision methods to estimate facial emotion patterns and stress-related visual signals.

These different signals are integrated using a weighted scoring model to generate an overall wellness score, confidence level, risk summary, and clinical flags. When the assessment indicates elevated concern, the system provides personalized recommendations, crisis-support guidance where necessary, and access to nearby mental health resources discovered via Google Places API or OpenStreetMap. The system stores session data such as questionnaire scores, transcript summaries, audio-video analysis results, and generated reports. By combining rule-based PHQ-9 scoring, Hugging Face transformer emotion models, OpenAI Whisper speech recognition, Deep Face facial analysis, and optional Gemini LLM synthesis, the project offers a practical real-time mental health screening and support platform that helps users reflect on their condition and seek timely professional assistance.

Keywords

Mental Health Monitoring, Multimodal Learning, Artificial Intelligence, Emotion Recognition, Real-Time Analysis

Downloads

References

1. A. Radford et al., "Robust speech recognition via large-scale weak supervision," in Proc. ICML, pp. 28492-28518, 2023. [Google Scholar] [Crossref]

2. S. I. Serengil and A. Ozpinar, "LightFace: A hybrid deep face recognition framework," in Proc. IEEE INISTA, pp. 1-5, 2020. [Google Scholar] [Crossref]

3. X. Yang, Y. Lenci, and B. W. Schuller, "Emotion recognition from speech with multi-task learning," Neurocomputing, vol. 391, pp. 279-288, 2020. [Google Scholar] [Crossref]

4. S. I. Serengil and A. Ozpinar, "LightFace: A hybrid deep face recognition framework," in [Google Scholar] [Crossref]

5. Proc. IEEE INISTA, pp. 1-5, 2020. [Google Scholar] [Crossref]

6. M. Jiang, P. Yang, and X. Bhanu, "Multi-modal depression detection using audio, visual, and text data," in Proc. IEEE ACII, pp. 1-7, 2019. [Google Scholar] [Crossref]

Metrics

Views & Downloads

Similar Articles