Bridging Learning and Passivity: Integral Reinforcement Learning with SPR Transformation for Multi-Agent Coordination

Authors

Jie Ren

Electrical & Computer Engineering, Northeastern University, Boston, MA (United States of America)

Bahram Shafai

Electrical & Computer Engineering, Northeastern University, Boston, MA (United States of America)

Article Information

DOI: 10.51244/IJRSI.2026.1307000205

Subject Category: Computer Science

Volume/Issue: 13/7 | Page No: 2748-2767

Publication Timeline

Submitted: 2026-07-22

Accepted: 2026-07-27

Published: 2026-08-07

Abstract

This paper presents a framework that enables multi-agent systems to learn controllers independently through integral reinforcement learning while guaranteeing modular stability when connected. The framework addresses a gap: existing modular control methods require analytical models and cannot handle learned controllers, while multi-agent learning methods require network coupling during training. The proposed approach combines Integral Reinforcement Learning for independent training with a Strictly Positive Real transformation via a Proportional-Integral observer for post-learning certification. We establish necessary and sufficient conditions for the PI observer construction and prove that SPR-transformed agents achieve output synchronization, regulation to zero, and global asymptotic stability under strongly connected graphs. Unlike standard passivity-based coordination that synchronizes to arbitrary constants, the positive definiteness condition required by strict positive realness ensures regulation to the origin. We derive a performance bound showing that the network convergence rate and the stability margin depend on the graph topology, the agent dynamics, and the controller quality. The framework trades local optimality for modular guarantees: controllers optimized for isolated operation become conservative when connected, but the performance bound quantifies this cost, showing that better local controllers reduce the penalty. Simulations on 10 heterogeneous agents validate the framework across 800 Monte Carlo trials, demonstrating that IRL achieves statistical equivalence to optimal LQR despite 20% model uncertainty and unstable agents.
Department of Electrical Engineering, Northeastern University, Boston, MA 02115, USA

Keywords

Cooperative control, multi-agent systems, reinforcement learning, passivity, optimal control

Downloads

References

1. Etika Agarwal, Shuvomoy Das Gupta, Necmiye Ozay, and Petros Voulgaris. Distributed synthesis of local controllers for networked systems with arbitrary interconnection topologies. IEEE Transactions on Automatic Control, 66 (2): 683–698, 2021. doi: 10.1109/TAC.2020.3028593. [Google Scholar] [Crossref]

2. Murat Arcak. Passivity as a design tool for group coordination. IEEE Transactions on Automatic Control, 52 (8): 1380–1390, 2007. doi: 10.1109/TAC.2007.902733. [Google Scholar] [Crossref]

3. S. Beale and B. Shafai. Robust control system design with a proportional integral observer. International Journal of Control, 50 (1): 97–111, 1989. [Google Scholar] [Crossref]

4. W. Brewer. Kronecker products and matrix calculus in system theory. IEEE Transactions on Circuits and Systems, 25 (9): 772–781, 1978. [Google Scholar] [Crossref]

5. Sven Gronauer and Klaus Diepold. Multi-agent deep reinforcement learning: a survey. Artificial Intelligence Review, 55 (2): 895–943, 2022. doi: 10.1007/s10462-021-09996-w. [Google Scholar] [Crossref]

6. Changchun Hua, Kuo Liu, and Xinping Guan. Semi-global/global output consensus for nonlinear multiagent systems with time delays. Automatica, 103: 480–489, 2019. doi: 10.1016/j.automatica.2019.02.022. [Google Scholar] [Crossref]

7. Yutao Jiang, Weinan Gao, Jing Wu, Tianyou Chai, and Frank L. Lewis. Reinforcement learning and cooperative H_∞ output regulation of linear continuous-time multi-agent systems. Automatica, 150: 110581, 2023. doi: 10.1016/j.automatica.2022.10634. [Google Scholar] [Crossref]

8. Qing Jiao, Hamidreza Modares, Shengyuan Xu, Frank L. Lewis, and Kyriakos G. Vamvoudakis. Multi-agent zero-sum differential graphical games for disturbance rejection in distributed control. Automatica, 69: 24–34, 2016. doi: 10.1016/j.automatica.2016.02.002. [Google Scholar] [Crossref]

9. Ruiyang Jin, Zaiwei Chen, Yiheng Lin, Jie Song, and Adam Wierman. Approximate global convergence of independent learning in multi-agent systems. In Proceedings of the 28th International Conference on Artificial Intelligence and Statistics (AISTATS), volume 258 of Proceedings of Machine Learning Research, pages 2818–2826. PMLR, 2025. [Google Scholar] [Crossref]

10. Yu Kawano, Krishna Chaitanya Kosaraju, and Jacquelien M. A. Scherpen. Krasovskii and shifted passivity-based control. IEEE Transactions on Automatic Control, 66 (10): 4926–4932, 2021. doi: 10.1109/TAC.2020.3040252. [Google Scholar] [Crossref]

11. Yu Kawano, Alessio Moreschini, and Michele Cucuzzella. Krasovskii passivity for sampled-data stabilization and output consensus. IEEE Transactions on Automatic Control, 70 (2): 1038–1053, 2025. doi: 10.1109/TAC.2024.3438749. [Google Scholar] [Crossref]

12. Hassan K. Khalil. Nonlinear Control. Pearson, 3rd edition, 2015. [Google Scholar] [Crossref]

13. Hyungbo Kim, Hyungbo Shim, Juhoon Back, and Jin Heon Seo. Consensus of output-coupled linear multi-agent systems under fast switching network: Averaging approach. Automatica, 49 (1): 267–272, 2013. doi: 10.1016/j.automatica.2012.09.025. [Google Scholar] [Crossref]

14. Jin Gyu Lee and Hyungbo Shim. A tool for analysis and synthesis of heterogeneous multi-agent systems under rank-deficient coupling. Automatica, 120: 109118, 2020. doi: 10.1016/j.automatica.2020.109118. [Google Scholar] [Crossref]

15. Frank L. Lewis, Hongwei Zhang, Kristian Hengster-Movric, and Abhijit Das. Cooperative Control of Multi-Agent Systems: Optimal and Adaptive Design Approaches. Communications and Control Engineering. Springer-Verlag London, 2014. doi: 10.1007/978-1-4471-5574-4. [Google Scholar] [Crossref]

16. Jianguo Li, Hamidreza Modares, Tianyou Chai, Frank L. Lewis, and Lihua Xie. Off-policy reinforcement learning for synchronization in multiagent graphical games. IEEE Transactions on Neural Networks and Learning Systems, 28 (10): 2434–2445, 2017. doi: 10.1109/TNNLS.2016.2609500. [Google Scholar] [Crossref]

17. Ryan Lowe, Yi Wu, Aviv Tamar, Jean Harb, Pieter Abbeel, and Igor Mordatch. Multi-agent actor-critic for mixed cooperative-competitive environments. In Advances in Neural Information Processing Systems 30 (NeurIPS 2017), pages 6379–6390, 2017. [Google Scholar] [Crossref]

18. Reza Olfati-Saber, J. Alex Fax, and Richard M. Murray. Consensus and cooperation in networked multi-agent systems. Proceedings of the IEEE, 95 (1): 215–233, 2007. doi: 10.1109/JPROC.2006.887293. [Google Scholar] [Crossref]

19. Afshin Oroojlooy and Davood Hajinezhad. A review of cooperative multi-agent deep reinforcement learning. Applied Intelligence, 53 (11): 13677–13722, 2023. doi: 10.1007/s10489-022-04105-y. [Google Scholar] [Crossref]

20. Zhihua Qu and Marwan A. Simaan. Modularized design for cooperative control and plug-and-play operation of networked heterogeneous systems. Automatica, 50 (9): 2405–2414, 2014. doi: 10.1016/j.automatica.2014.07.003. [Google Scholar] [Crossref]

21. Tabish Rashid, Mikayel Samvelyan, Christian Schroeder de Witt, Gregory Farquhar, Jakob Foerster, and Shimon Whiteson. Monotonic value function factorisation for deep multi-agent reinforcement learning. Journal of Machine Learning Research, 21 (178): 1–51, 2020. [Google Scholar] [Crossref]

22. Jie Ren and Bahram Shafai. Transformation of linear systems into strictly positive real systems with a proportional-integral observer. In The 2025 Modeling, Estimation and Control Conference, 2025. [Google Scholar] [Crossref]

23. Wei Ren and Randal W. Beard. Consensus seeking in multiagent systems under dynamically changing interaction topologies. IEEE Transactions on Automatic Control, 50 (5): 655–661, 2005. doi: 10.1109/TAC.2005.846556. [Google Scholar] [Crossref]

24. Stefano Riverso, Marcello Farina, and Giancarlo Ferrari-Trecate. Plug-and-Play Model Predictive Control based on robust control invariant sets. Automatica, 50 (8): 2179–2186, 2014. doi: 10.1016/j.automatica.2014.06.005. [Google Scholar] [Crossref]

25. Georg S. Seyboth, Wei Ren, and Frank Allgöwer. Cooperative control of linear multi-agent systems via distributed output regulation and transient synchronization. Automatica, 68: 132–139, 2016. doi: 10.1016/j.automatica.2016.01.068. [Google Scholar] [Crossref]

26. B. Shafai and M. Saif. Proportional-integral observer in robust control, fault detection, and decentralized control of dynamic systems. In Robust Control and Fault Detection, volume 27, pages 13–43. Springer International Publishing, Cham, 2015. [Google Scholar] [Crossref]

27. Takashi Tanaka and Cedric Langbort. Retrofit control with approximate environment modeling. Automatica, 108: 108488, 2019. doi: 10.1016/j.automatica.2019.06.031. [Google Scholar] [Crossref]

28. Farzaneh Tatari, Mohammad Bagher Naghibi-Sistani, and Kyriakos G. Vamvoudakis. Distributed learning algorithm for non-linear differential graphical games. Transactions of the Institute of Measurement and Control, 39 (2): 173–182, 2017. doi: 10.1177/0142331215603791. [Google Scholar] [Crossref]

29. Kyriakos G. Vamvoudakis and Frank L. Lewis. Multi-player non-zero-sum games: Online adaptive learning solution of coupled Hamilton–Jacobi equations. Automatica, 47 (8): 1556–1569, 2011. doi: 10.1016/j.automatica.2011.03.005. [Google Scholar] [Crossref]

30. Kyriakos G. Vamvoudakis, Frank L. Lewis, and Greg R. Hudas. Multi-agent differential graphical games: Online adaptive learning solution for synchronization with optimality. Automatica, 48 (8): 1598–1611, 2012. doi: 10.1016/j.automatica.2012.03.005. [Google Scholar] [Crossref]

31. Draguna Vrabie, Octavian Pastravanu, Murad Abu-Khalaf, and Frank L. Lewis. Adaptive optimal control for continuous-time linear systems based on policy iteration. Automatica, 45 (2): 477–484, 2009. doi: 10.1016/j.automatica.2008.08.017. [Google Scholar] [Crossref]

32. Peter Wieland, Rodolphe Sepulchre, and Frank Allgöwer. An internal model principle is necessary and sufficient for linear output synchronization. Automatica, 47 (5): 1068–1074, 2011. doi: 10.1016/j.automatica.2011.01.081. [Google Scholar] [Crossref]

33. Chao Yu, Akash Velu, Eugene Vinitsky, Jiaxuan Gao, Yu Wang, Alexandre M. Bayen, and Yi Wu. The surprising effectiveness of PPO in cooperative multi-agent games. In Advances in Neural Information Processing Systems 35 (NeurIPS 2022), 2022. [Google Scholar] [Crossref]

34. Huaguang Zhang, He Jiang, Yanhong Luo, and Geyang Xiao. Data-driven optimal consensus control for discrete-time multi-agent systems with unknown dynamics using reinforcement learning method. IEEE Transactions on Industrial Electronics, 64 (5): 4091–4100, 2017. doi: 10.1109/TIE.2016.2542134. [Google Scholar] [Crossref]

Metrics

Views & Downloads

Similar Articles