系统工程与电子技术 ›› 2026, Vol. 48 ›› Issue (9): 3126-3135.doi: 10.12305/j.issn.1001-506X.2026.09.25

• 系统工程 • 上一篇    

逻辑链构建下基于深度强化学习的多智能体协同算法优化

张杰1,2(), 朱峰1, 王超1, 李东1, 陈志群1, 吕天启1   

  1. 1. 中国电子科技集团公司第二十八研究所,江苏 南京 210000
    2. 南京大学计算机学院,江苏 南京 210023
  • 收稿日期:2024-11-20 修回日期:2025-04-14 接受日期:2026-03-18 出版日期:2026-03-20 发布日期:2026-03-20
  • 通讯作者: 张杰 E-mail:guyuexiao95@gmail.com
  • 作者简介:朱 峰(1985—),男,研究员,高级工程师,硕士,主要研究方向为指挥控制系统
    王 超(1988—),男,研究员,高级工程师,博士,主要研究方向为指挥控制系统
    李 东(1992—),男,工程师,博士,主要研究方向为指挥控制系统
    陈志群(1991—),男,工程师,博士,主要研究方向为指挥控制系统
    吕天启(1992—),男,工程师,博士,主要研究方向为指挥控制系统

Optimization of multi-agent cooperation algorithm based on deep reinforcement learning for logical chain

Jie Zhang1,2(), Feng Zhu1, Chao Wang1, Dong Li1, Zhiqun Chen1, Tianqi Lyu1   

  1. 1. The 28th Research Institute of China Electronics Technology Group Corporation,Nanjing 210000,China
    2. School of Computer Science,Nanjing University,Nanjing 210023,China
  • Received:2024-11-20 Revised:2025-04-14 Accepted:2026-03-18 Online:2026-03-20 Published:2026-03-20
  • Contact: Jie Zhang E-mail:guyuexiao95@gmail.com

摘要:

在现代战争中,快速、精确的场景决策和有效的资源调度是提高作战效率的关键,因此如何构建高效的逻辑链成为核心问题之一。基于逻辑链构建任务提出一种基于深度强化学习的多智能体协同算法,将场景环境建模成多智能体系统,涵盖侦察、决策、打击和后勤等环节,系统内智能体分别负责不同的任务分工,并基于 Actor-Critic 训练。通过协同学习机制,各智能体能够动态调整策略,实现全局目标的优化。实验结果表明,与对比算法相比,所提出方法在任务成功率、资源利用效率及决策稳定性等方面均取得更优表现,其中任务成功率平均提升约 4%~8%,能够有效提升复杂场景环境下逻辑链构建与协同决策的整体效率。

关键词: 深度强化学习, 多智能体, Actor-Critic, 逻辑链, 协同算法优化

Abstract:

In modern warfare, rapid and precise battlefield decision-making, along with efficient resource allocation, is critical to enhancing combat effectiveness. Consequently, constructing an efficient kill chain has become one of the core challenges. This paper proposes a multi-agent collaboration algorithm based on deep reinforcement learning for kill chain construction tasks. The battlefield environment is modeled as a multi-agent system, encompassing reconnaissance, decision-making, strikes, and logistics. Each agent within the system is responsible for different task divisions and is trained based on the Actor-Critic framework. Through a cooperative learning mechanism, agents can dynamically adjust their strategies to optimize the global objectives. Experimental results demonstrate that, compared with baseline algorithms, the proposed algorithm achieves superior performance in terms of task success rate, resource utilization efficiency, and decision stability. In particular, the task success rate is improved by approximately 4%–8%, indicating that the proposed algorithm can effectively enhance the efficiency of kill-chain construction and collaborative decision-making in complex battlefield environments.

Key words: deep reinforcement learning, multi-agent, Actor-Critic, logical chain, cooperative algorithm optimization

中图分类号: