系统工程与电子技术 ›› 2026, Vol. 48 ›› Issue (10): 3531-3540.doi: 10.12305/j.issn.1001-506X.2026.10.23

• 系统工程 • 上一篇    

基于深度分层强化学习的防空作战武器目标分配算法

曹波1, 邢清华2, 刘家义3, 李龙跃2   

  1. 1. 空军工程大学研究生院,陕西 西安 710051
    2. 空军工程大学防空反导学院,陕西 西安 710051
    3. 国防大学联合作战学院,河北 石家庄 050084
  • 收稿日期:2025-07-14 出版日期:2026-10-25 发布日期:2026-09-30
  • 通讯作者: 李龙跃
  • 作者简介:曹 波(1998—),男,博士研究生,主要研究方向为深度学习与智能优化算法、态势认知
    邢清华(1966—),女,教授,博士,主要研究方向为智能作战指挥决策优化理论与方法
    刘家义(1995—),男,讲师,博士,主要研究方向为深度学习与智能优化算法
  • 基金资助:
    国家自然科学基金(72071209)资助课题

Air defense operations weapon target assignment algorithm based on deep hierarchical reinforcement learning

Bo Cao1, Qinghua Xing2, Jiayi Liu3, Longyue Li2   

  1. 1. Graduate School,Air Force Engineering University,Xi’an 710051,China
    2. Air Defense and Antimissile School,Air Force Engineering University,Xi’an 710051,China
    3. Joint Operations College,National Defense University,Shijiazhuang 050084,China
  • Received:2025-07-14 Online:2026-10-25 Published:2026-09-30
  • Contact: Longyue Li

摘要:

针对传统求解方法在大规模作战场景中面临维数灾难和实时性挑战等问题,提出一种基于深度分层强化学习的防空作战武器目标分配算法。首先构建以最小化蓝方预期生存威胁和红方弹药消耗为目标的动态分配模型,之后将目标分配任务分解为战略层任务匹配与战术层目标拦截,分别采用近端策略优化(proximal policy optimization, PPO)算法和融合注意力机制的多智能体PPO算法求解,并引入三阶段课程学习机制提升模型泛化能力。实验结果表明,所提方法在效费比和收敛效率上优于传统强化学习算法和启发式算法,为提升防空作战决策效率提供了有效支撑。

关键词: 防空作战目标分配, 深度分层强化学习, 近端策略优化算法, 课程学习

Abstract:

A deep hierarchical reinforcement learning based air defense weapon target assignment algorithm is proposed to address the challenges of dimensionality disaster and real-time performance faced by traditional solving methods in large-scale combat scenarios. Firstly, a dynamic assignment model is constructed with the goal of minimizing the expected survival threat of the blue side and the ammunition consumption of the red side. Then, the target assignment task is decomposed into strategic level task matching and tactical level target interception, which are solved using the proximal policy optimization (PPO) algorithm and the multi-agent PPO algorithm incorporating attention mechanism, respectively. A three-stage curriculum learning mechanism is introduced to improve the generalization ability of the model. The experimental results show that the proposed method is superior to traditional reinforcement learning algorithms and heuristic algorithms in terms of cost-effectiveness and convergence efficiency, providing effective support for improving the decision-making efficiency of air defense operations.

Key words: target assignment for air defense operations, deep hierarchical reinforcement learning (DHRL), proximal policy optimization (PPO) algorithm, curriculum learning

中图分类号: