Multi-Model Large Language Models for Human-Robot Interaction
Authors
Assistant Professor, Axis Institute of Higher Education (Uttar Pradesh)
Assistant Professor, Axis Institute of Higher Education (Uttar Pradesh)
Assistant Professor, Axis Institute of Higher Education (Uttar Pradesh)
Assistant Professor, Axis Institute of Higher Education (Uttar Pradesh)
Article Information
DOI: 10.47772/IJRISS.2026.100900212
Subject Category: Language
Volume/Issue: 10/9 | Page No: 3084-3094
Publication Timeline
Submitted: 2026-09-09
Accepted: 2026-09-14
Published: 2026-10-06
Abstract
Human–robot interaction (HRI) becomes challenging when robots must interpret human instructions expressed through multiple forms of communication. People often combine speech, pointing gestures, visual cues, and contextual information, and these signals may be incomplete, ambiguous, or inconsistent. Multimodal large language models (MLLMs) can combine such inputs to support more natural interaction, but unreliable interpretations and hallucinated information may still lead to incorrect robotic decisions. This paper proposes a reliability-focused multimodal framework for MLLM-based HRI that integrates voice, pointing gestures, visual perception, and contextual cues for human intention and object understanding. The framework processes information from each modality and combines the resulting evidence within the MLLM reasoning stage. It then examines whether the modalities support the same interpretation, identifies conflicts, estimates uncertainty, and verifies the predicted interpretation against available visual and sensor evidence. When the evidence provides sufficient support for a safe decision, the system selects the appropriate response. When the information is ambiguous, contradictory, or poorly supported, the system asks the user for clarification instead of making an unsupported assumption. Contextual and emotional cues are used as additional evidence to improve interpretation rather than as definitive measures of a person's internal state. The proposed framework is evaluated through controlled HRI scenarios involving ambiguous instructions, conflicting speech and gestures, contextual changes, and hallucination-related cases. Evaluation considers intention accuracy, conflict-detection accuracy, hallucination rate, clarification effectiveness, response time, task-completion rate, and unsafe-action rate. By combining multimodal reasoning with uncertainty estimation, verification, and clarification, the proposed framework aims to improve the reliability, safety, and trustworthiness of MLLM-based human–robot interaction.
Keywords
Human–Robot Interaction, Multimodal Large Language Models, Human Intention Understanding, Multimodal Fusion, Conflict Resolution
Downloads
References
1. D. Driesset al., “PaLM-E: An Embodied Multimodal Language Model,” Proceedings of the 40th International Conference on Machine Learning, vol. 202, pp. 8469–8488, 2023. [Google Scholar] [Crossref]
2. B. Ichteret al., “Do As I Can, Not As I Say: Grounding Language in Robotic Affordances,” Proceedings of the 6th Conference on Robot Learning, vol. 205, pp. 287–318, 2023. [Google Scholar] [Crossref]
3. W. Huang et al., “Inner Monologue: Embodied Reasoning through Planning with Language Models,” Proceedings of the 6th Conference on Robot Learning, vol. 205, pp. 1769–1782, 2023. [Google Scholar] [Crossref]
4. B. Zitkovichet al., “RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control,”Proceedings of the 7th Conference on Robot Learning, vol. 229, pp. 2165–2183, 2023. [Google Scholar] [Crossref]
5. S. Trick, D. Koert, J. Peters, and C. Rothkopf, “Multimodal Uncertainty Reduction for Intention Recognition in Human-Robot Interaction,” arXiv preprint arXiv:1907.02426, 2019. [Google Scholar] [Crossref]
6. X. Zhao, H. Li, T. Miao, X. Zhu, Z. Wei, and A. Song, “Learning Multimodal Confidence for Intention Recognition in Human-Robot Interaction,” arXiv preprint arXiv:2405.14116, 2024. [Google Scholar] [Crossref]
7. S. Duncan, S. Alambeigi, and J. Pryor, “A Survey of Multimodal Perception Methods for Human–Robot Interaction in Social Environments,” ACM Transactions on Human-Robot Interaction, vol. 13, no. 4, 2024. [Google Scholar] [Crossref]
8. Y. Lai et al., “Natural Multimodal Fusion-Based Human–Robot Interaction: Application With Voice and Deictic Posture via Large Language Model,” IEEE Robotics and Automation Magazine, vol. 33, no. 2, pp. 29–38, 2026. [Google Scholar] [Crossref]
9. M. Abokiet al., “Multimodal Human–Robot Interaction Using Human Pose Estimation and Local Large Language Models,” Advanced Robotics Research, 2026, doi: 10.1002/adrr.202500175. [Google Scholar] [Crossref]
10. S. Yeon, M. Lee, J. Bae, and S. Lee, “Resolving Ambiguity in Pointing Gestures Using Contextual Reasoning from Large Language Models,” Computer Modeling in Engineering & Sciences, vol. 147, no. 1, 2026, doi: 10.32604/cmes.2026.079954. [Google Scholar] [Crossref]
11. C. Li et al., “Do MLLMs Understand Pointing? Benchmarking and Enhancing Referential Reasoning in Egocentric Vision,” Findings of the Association for Computational Linguistics: ACL 2026, 2026. [Google Scholar] [Crossref]
12. T. Schreiter et al., “Multimodal Interaction and Intention Communication for Industrial Robots,” arXiv preprint arXiv:2502.17971, 2025. [Google Scholar] [Crossref]
13. G. A. Abbo, S. Lenaerts, and T. Belpaeme, “Multimodal Large Language Models for Real-Time Situated Reasoning,” Companion Proceedings of the 21st ACM/IEEE International Conference on Human-Robot Interaction, pp. 1161–1163, 2026. [Google Scholar] [Crossref]
14. S. Spezialetti, G. Placidi, and M. Rossi, “Emotion Recognition for Human-Robot Interaction: Recent Advances and Future Perspectives,” Frontiers in Robotics and AI, vol. 7, 532279, 2020. [Google Scholar] [Crossref]
15. C. Bai et al., “Hallucination of Multimodal Large Language Models: A Survey,” arXiv preprint arXiv:2404.18930, 2024. [Google Scholar] [Crossref]
Metrics
Views & Downloads
Similar Articles
- Evaluating the Impacts of Mind Mapping Strategy on Developing EFL Students’ Critical Reading Skills
- Significance of Reading Instructions for Language Improvement in Children with Down Syndrome
- Prenasalised Consonants in Liangmai
- Metadiscourse Matters: Definitions, Models, and Advantages for ESL/ EFL Writing
- Blank Minds and Stuck Voices: Understanding and Addressing Cognitive Anxiety in High-Stakes ESL Speaking Tests