系统工程与电子技术 ›› 2026, Vol. 48 ›› Issue (9): 3089-3103.doi: 10.12305/j.issn.1001-506X.2026.09.22

• 系统工程 • 上一篇    

基于深度强化学习的多智能体集中式作战行动指挥决策方法

赵立阳(), 吴金平, 杨静, 张瑞   

  1. 海军潜艇学院,山东 青岛 266199
  • 收稿日期:2025-09-09 修回日期:2025-11-23 接受日期:2025-12-15 出版日期:2026-06-09 发布日期:2026-06-09
  • 通讯作者: 赵立阳 E-mail:kiumilitay@163.com
  • 作者简介:吴金平(1976—),男,研究员,博士,主要研究方向为系统建模与仿真、军事智能决策
    杨 静(1989—),女,副研究员,博士,主要研究方向为军事智能决策、深度强化学习
    张 瑞(2001—),男,博士研究生,主要研究方向为军事智能决策、深度强化学习
  • 基金资助:
    翱翔实验室重点项目(2023-CXPT-LC-003)资助课题

Centralized command decision-making method for multi-agent combat operations based on deep reinforcement learning

Liyang Zhao(), Jinping Wu, Jing Yang, Rui Zhang   

  1. Naval Submarine Academy,Qingdao 266199,China
  • Received:2025-09-09 Revised:2025-11-23 Accepted:2025-12-15 Online:2026-06-09 Published:2026-06-09
  • Contact: Liyang Zhao E-mail:kiumilitay@163.com

摘要:

针对复杂动态对抗环境下多智能体协同作战行动决策一致性差的问题,提出一种基于深度强化学习的多智能体集中式作战行动指挥决策方法。首先,为更好地控制多智能体产生作战意图协调一致的作战行动,对多智能体协同作战过程进行了马尔可夫决策过程建模和形式化描述,构建了区分标量、实体信息的状态空间和离散复合的动作空间,提出基于任务目标分解和贡献度分配的奖励重塑方法,设计与之相适配的基于多特征信息编码、非完全信息推理和多头复合动作解码的集中式作战决策智能体策略网络模型;其次,为提高计算资源的利用率和样本效率以及确保网络训练过程的稳定性,提出基于行为者−学习者的分布式采样训练框架和使用带有策略熵的SARD-PPO算法对模型进行训练;最后,在高保真度的仿真推演平台上进行仿真实验,验证了技术路径的可行性和有效性,通过调用决策智能体模型进行前向推演,实现了对抗过程的复盘分析和总结。

关键词: 对抗博弈, 集中式决策, 协同作战, 智能决策, 深度强化学习

Abstract:

To address the issue of poor decision-making consistency in multi-agent cooperative operations within complex dynamic adversarial environments, a centralized command decision-making method for multi-agent combat operations based on deep reinforcement learning. Firstly, to better control multiple agents in generating combat actions with coordinated intentions, the multi-agent cooperative combat process is modeled and formally described as a Markov decision process. A state space distinguishing between scalar and entity information, along with a discrete composite action space, is constructed. A reward reshaping method based on task objective decomposition and contribution allocation is proposed. A centralized operational decision-making agent policy network model is designed, incorporating multi-feature information encoding, incomplete information reasoning, and multi-head composite action decoding. Secondly, to improve computational resource utilization and sample efficiency while ensuring the stability of the network training process, an distributed sampling training framework based on Actor-Learner is proposed. The model is trained using the SARD-PPO algorithm with policy entropy. Finally, simulation experiments conducted on a high-fidelity simulation platform verify the feasibility and effectiveness of the technical approach. By deploying the decision-making agent model for forward simulation, replay analysis and summary of the adversarial process are achieved.

Key words: adversarial games, centralized decision-making, collaborative operations, intelligent decision-making, deep reinforcement learning

中图分类号: