Systems Engineering and Electronics ›› 2026, Vol. 48 ›› Issue (5): 1670-1681.doi: 10.12305/j.issn.1001-506X.2026.05.23

• Systems Engineering • Previous Articles     Next Articles

Robust multi-agent cooperative confrontation policy offline reinforcement learning

Huaqing ZHANG1, Xiaofei ZHANG2, Mingrui HAO1, Jixiang JIANG1,*, Shan LI3   

  1. 1. National Key Laboratory of Complex System Control and Intelligent Agent Cooperation,Beijing Institute of Mechanical and Electrical Engineering,Beijing 100074,China
    2. Beijing Institute of Computer Technology and Applications, Beijing 100854,China
    3. School of Mathematics and Statistics,Hainan University,Haikou 570228,China
  • Received:2025-01-14 Online:2026-05-27 Published:2026-05-27
  • Contact: Jixiang JIANG

Abstract:

To address the problem of offline learning of multi-agent confrontation policies in dynamic scenarios, a robust multi-agent policy offline reinforcement learning (RMA-offlineRL) method is proposed, aiming to reduce the impact of dataset quality on offline policy learning. In the policy improvement of RMA-offlineRL, the offline policy gradient is calculated by performing a Box-Cox transformation on the log-probabilities of historical state-action pairs sampled from the dataset under the current policy. It can not only limit extrapolation errors but also improve the policy using low-quality datasets containing large amounts of random behaviors. In addition, in the policy evaluation of RMA-offlineRL, a stable policy evaluation algorithm is designed for multi-step offline interaction datas. To verify the effectiveness and advantages of the proposed method, comprehensive validation is conducted on offline datasets collected using MiaoSuan wargame system and benchmark environments. Results show that the proposed method can learn cooperative confrontation policies efficiently from low-quality datasets containing large amounts of random behaviors or with a low state-action coverage index.

Key words: multi-agent, cooperative confrontation, offline reinforcement learning, confrontation scenarios, random behavior

CLC Number: 

[an error occurred while processing this directive]