Multi-Model Large Language Models for Human-Robot Interaction

Authors

Amarnath Awasthi

Assistant Professor, Axis Institute of Higher Education (Uttar Pradesh)

Abhinav Singh

Assistant Professor, Axis Institute of Higher Education (Uttar Pradesh)

Aastha Gupta

Assistant Professor, Axis Institute of Higher Education (Uttar Pradesh)

Aditya Prajapati

Assistant Professor, Axis Institute of Higher Education (Uttar Pradesh)

Article Information

DOI: 10.47772/IJRISS.2026.100900212

Subject Category: Language

Volume/Issue: 10/9 | Page No: 3084-3094

Publication Timeline

Submitted: 2026-09-09

Accepted: 2026-09-14

Published: 2026-10-06

Abstract

Human–robot interaction (HRI) becomes challenging when robots must interpret human instructions expressed through multiple forms of communication. People often combine speech, pointing gestures, visual cues, and contextual information, and these signals may be incomplete, ambiguous, or inconsistent. Multimodal large language models (MLLMs) can combine such inputs to support more natural interaction, but unreliable interpretations and hallucinated information may still lead to incorrect robotic decisions. This paper proposes a reliability-focused multimodal framework for MLLM-based HRI that integrates voice, pointing gestures, visual perception, and contextual cues for human intention and object understanding. The framework processes information from each modality and combines the resulting evidence within the MLLM reasoning stage. It then examines whether the modalities support the same interpretation, identifies conflicts, estimates uncertainty, and verifies the predicted interpretation against available visual and sensor evidence. When the evidence provides sufficient support for a safe decision, the system selects the appropriate response. When the information is ambiguous, contradictory, or poorly supported, the system asks the user for clarification instead of making an unsupported assumption. Contextual and emotional cues are used as additional evidence to improve interpretation rather than as definitive measures of a person's internal state. The proposed framework is evaluated through controlled HRI scenarios involving ambiguous instructions, conflicting speech and gestures, contextual changes, and hallucination-related cases. Evaluation considers intention accuracy, conflict-detection accuracy, hallucination rate, clarification effectiveness, response time, task-completion rate, and unsafe-action rate. By combining multimodal reasoning with uncertainty estimation, verification, and clarification, the proposed framework aims to improve the reliability, safety, and trustworthiness of MLLM-based human–robot interaction.

Keywords

Human–Robot Interaction, Multimodal Large Language Models, Human Intention Understanding, Multimodal Fusion, Conflict Resolution

Downloads

References

1. D. Driesset al., “PaLM-E: An Embodied Multimodal Language Model,” Proceedings of the 40th International Conference on Machine Learning, vol. 202, pp. 8469–8488, 2023. [Google Scholar] [Crossref]

2. B. Ichteret al., “Do As I Can, Not As I Say: Grounding Language in Robotic Affordances,” Proceedings of the 6th Conference on Robot Learning, vol. 205, pp. 287–318, 2023. [Google Scholar] [Crossref]

3. W. Huang et al., “Inner Monologue: Embodied Reasoning through Planning with Language Models,” Proceedings of the 6th Conference on Robot Learning, vol. 205, pp. 1769–1782, 2023. [Google Scholar] [Crossref]

4. B. Zitkovichet al., “RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control,”Proceedings of the 7th Conference on Robot Learning, vol. 229, pp. 2165–2183, 2023. [Google Scholar] [Crossref]

5. S. Trick, D. Koert, J. Peters, and C. Rothkopf, “Multimodal Uncertainty Reduction for Intention Recognition in Human-Robot Interaction,” arXiv preprint arXiv:1907.02426, 2019. [Google Scholar] [Crossref]

6. X. Zhao, H. Li, T. Miao, X. Zhu, Z. Wei, and A. Song, “Learning Multimodal Confidence for Intention Recognition in Human-Robot Interaction,” arXiv preprint arXiv:2405.14116, 2024. [Google Scholar] [Crossref]

7. S. Duncan, S. Alambeigi, and J. Pryor, “A Survey of Multimodal Perception Methods for Human–Robot Interaction in Social Environments,” ACM Transactions on Human-Robot Interaction, vol. 13, no. 4, 2024. [Google Scholar] [Crossref]

8. Y. Lai et al., “Natural Multimodal Fusion-Based Human–Robot Interaction: Application With Voice and Deictic Posture via Large Language Model,” IEEE Robotics and Automation Magazine, vol. 33, no. 2, pp. 29–38, 2026. [Google Scholar] [Crossref]

9. M. Abokiet al., “Multimodal Human–Robot Interaction Using Human Pose Estimation and Local Large Language Models,” Advanced Robotics Research, 2026, doi: 10.1002/adrr.202500175. [Google Scholar] [Crossref]

10. S. Yeon, M. Lee, J. Bae, and S. Lee, “Resolving Ambiguity in Pointing Gestures Using Contextual Reasoning from Large Language Models,” Computer Modeling in Engineering & Sciences, vol. 147, no. 1, 2026, doi: 10.32604/cmes.2026.079954. [Google Scholar] [Crossref]

11. C. Li et al., “Do MLLMs Understand Pointing? Benchmarking and Enhancing Referential Reasoning in Egocentric Vision,” Findings of the Association for Computational Linguistics: ACL 2026, 2026. [Google Scholar] [Crossref]

12. T. Schreiter et al., “Multimodal Interaction and Intention Communication for Industrial Robots,” arXiv preprint arXiv:2502.17971, 2025. [Google Scholar] [Crossref]

13. G. A. Abbo, S. Lenaerts, and T. Belpaeme, “Multimodal Large Language Models for Real-Time Situated Reasoning,” Companion Proceedings of the 21st ACM/IEEE International Conference on Human-Robot Interaction, pp. 1161–1163, 2026. [Google Scholar] [Crossref]

14. S. Spezialetti, G. Placidi, and M. Rossi, “Emotion Recognition for Human-Robot Interaction: Recent Advances and Future Perspectives,” Frontiers in Robotics and AI, vol. 7, 532279, 2020. [Google Scholar] [Crossref]

15. C. Bai et al., “Hallucination of Multimodal Large Language Models: A Survey,” arXiv preprint arXiv:2404.18930, 2024. [Google Scholar] [Crossref]

Metrics

Views & Downloads

Similar Articles