Interactive Teaching of Robotics Based on Multimodal Human Pose Estimation

Authors

Xianhua Li

School of Mechatronics Engineering, Anhui University of Science and Technology, Huainan 232001, Anhui, China (China)

Minhui Wang

School of Mechatronics Engineering, Anhui University of Science and Technology, Huainan 232001, Anhui, China (China)

Zisen Hua

School of Artificial Intelligence, Anhui University of Science and Technology, Huainan 232001, Anhui, China (China)

Junjie Mi

School of Mechatronics Engineering, Anhui University of Science and Technology, Huainan 232001, Anhui, China (China)

Long Li

School of Artificial Intelligence, Anhui University of Science and Technology, Huainan 232001, Anhui, China (China)

Zhipeng Yu

School of Artificial Intelligence, Anhui University of Science and Technology, Huainan 232001, Anhui, China (China)

Article Information

DOI: 10.47772/IJRISS.2026.100900089

Subject Category: Social science

Volume/Issue: 10/9 | Page No: 1348-1356

Publication Timeline

Submitted: 2026-09-20

Accepted: 2026-09-25

Published: 2026-09-30

Abstract

Traditional robotics courses often provide limited opportunities for engineering practice, making it difficult for students to apply theoretical knowledge to integrated engineering tasks. To address this issue, this study proposes a practical robotics teaching approach integrating multimodal vision and human pose estimation. An RGB-IR multimodal human pose dataset is constructed using an Intel RealSense D435 camera, and an RGB-IR dual-branch pose estimation network with a Cross-Modal Adaptive Gated Fusion (CM-AGF) mechanism is developed based on YOLOv8-Pose. The detected two-dimensional keypoints are combined with depth information for three-dimensional reconstruction and joint-angle calculation, and the resulting human motion parameters are mapped to robot control commands. Experimental results demonstrate the feasibility of integrating multimodal perception, pose estimation, three-dimensional motion reconstruction, and robot control into a unified practical teaching platform. The proposed approach provides students with opportunities to integrate computer vision, artificial intelligence, and robotics in engineering practice.

Keywords

robotics; multimodal vision; human pose estimation; teaching practice; human-robot interaction

Downloads

References

1. Lu, L., Dong, Q., Zhang, T., et al. (2023) Review of robot kinematics and motion planning algorithms, Research on Printing and Digital Media Technology, (5), 1–16. (in Chinese) [Google Scholar] [Crossref]

2. Xiong, Y., Tang, L., Ding, H., et al. (2008) Fundamentals of Robotics Technology, Huazhong University of Science and Technology Press, Wuhan, China. (in Chinese) [Google Scholar] [Crossref]

3. Yang, K., Yu, T., Wei, W., et al. (2025) Reform and practice of the blended teaching model for Robot Perception Technology under the background of emerging engineering education, Intelligent Manufacturing, (3), 124–128. (in Chinese) [Google Scholar] [Crossref]

4. Xu, G. and Wu, X. (2023) Application of a human pose estimation method based on the LSA-HRNet network in Tai Chi movements, Journal of South-Central Minzu University (Natural Science Edition), 42(6), 839–845. (in Chinese) [Google Scholar] [Crossref]

5. Upadhyay, A., Dhupar, B., Sharma, M., et al. (2024) LWIRPose: A novel LWIR thermal image dataset and benchmark, arXiv preprint, arXiv:2404.10212. [Google Scholar] [Crossref]

6. Tang, H., Zhang, X., Zhuo, S., et al. (2015) High resolution photography with an RGB-infrared camera, in 2015 IEEE International Conference on Computational Photography (ICCP), pp. 1–10. [Google Scholar] [Crossref]

7. Du, J., Zhou, H., Qian, K., et al. (2020) RGB-IR cross input and sub-pixel upsampling network for infrared image super-resolution, Sensors, 20(1), 281. [Google Scholar] [Crossref]

8. Chen, H., Feng, R., Wu, S., et al. (2023) 2D human pose estimation: A survey, Multimedia Systems, 29(5), 3115–3138. [Google Scholar] [Crossref]

9. Dong, C., Tang, Y. and Zhang, L. (2024) HDA-Pose: A real-time 2D human pose estimation method based on modified YOLOv8, Signal, Image and Video Processing, 18(8), 5823–5839. [Google Scholar] [Crossref]

10. Zhang, Y., Yan, W., Shan, T., et al. (2026) Research on passenger posture prediction in subway station based on lightweight improved YOLOv8-Pose, Electronics, 15(14), 3031. [Google Scholar] [Crossref]

Metrics

Views & Downloads

Similar Articles