系统工程与电子技术 ›› 2026, Vol. 48 ›› Issue (9): 3167-3178.doi: 10.12305/j.issn.1001-506X.2026.09.29

• 制导、导航与控制 • 上一篇    

深度强化学习驱动的多约束分数阶滑模制导律

张倬境, 盛永智(), 龚孝龙   

  1. 北京理工大学自动化学院,北京 100081
  • 收稿日期:2025-06-13 修回日期:2025-08-09 出版日期:2025-11-25 发布日期:2025-11-25
  • 通讯作者: 盛永智 E-mail:shengyongzhi@bit.edu.cn
  • 作者简介:张倬境(2001—),男,硕士研究生,主要研究方向为飞行器制导与控制、深度强化学习
    龚孝龙(2001—),男,硕士研究生,主要研究方向为飞行器制导与控制

Deep reinforcement learning-driven multi-constraint fractional order sliding mode guidance law

Zhuojing Zhang, Yongzhi Sheng(), Xiaolong Gong   

  1. School of Automation,Beijing Institute of Technology,Beijing 100081,China
  • Received:2025-06-13 Revised:2025-08-09 Online:2025-11-25 Published:2025-11-25
  • Contact: Yongzhi Sheng E-mail:shengyongzhi@bit.edu.cn

摘要:

针对高超声速拦截弹末端制导存在的攻击角、视场角等多约束难题,提出深度强化学习驱动的多约束分数阶滑模制导律。改进双延迟深度确定性策略梯度算法,引入多头注意力机制作为特征提取层,实现环境状态多尺度特征融合,利用长短期记忆网络增强时序特征理解,通过自适应探索与回合制训练优化框架提升算法收敛性与稳定性。基于弹目三维相对运动学模型,构建含攻击角约束的分数阶滑模基础制导律框架,通过改进强化学习算法动态优化制导参数,解决多约束下精确制导问题。仿真结果表明,所提制导律在满足多种角度约束条件下可显著提升制导速度与精度,验证了所提方法的正确性与有效性。

关键词: 制导律, 攻击角约束, 视场角约束, 分数阶滑模控制, 深度强化学习

Abstract:

A deep reinforcement learning-driven multi-constraint fractional order sliding mode guidance law is proposed to address the challenges of multi-constraint such as attack angle and field of view in the end guidance of hypersonic interception missiles. Dual delay depth deterministic strategy gradient algorithm is improved, multi-head attention mechanism is introduced as the feature extraction layer, multi-scale feature fusion of environmental states is achieved, long short-term memory network is utilized to enhance temporal feature understanding, and convergence and stability of the algorithm is improved through adaptive exploration and turn-based training optimization framework. Based on the three-dimensional relative kinematics model of the missile target, a fractional order sliding mode basic guidance law framework with attack angle constraints is constructed. By improving the reinforcement learning algorithm, the guidance parameters are dynamically optimized to solve the problem of precise guidance under multi-constraint. Simulation results show that the proposed guidance law can significantly improve the guidance speed and accuracy under various angle constraints. The correctness and effectiveness of the proposed method are verified.

Key words: guidance law, attack angle constraint, field of view constraint, fractional order sliding mode control, deep reinforcement learning

中图分类号: