系统工程与电子技术 ›› 2026, Vol. 48 ›› Issue (9): 3266-3273.doi: 10.12305/j.issn.1001-506X.2026.09.39

• 通信与网络 • 上一篇    

基于信息调节LSTM的抗干扰决策算法

刘锡国1,2(), 刘敏1,2, 毛忠阳1,2, 徐盟飞3, 马鹏3   

  1. 1. 海军航空大学,山东 烟台 264001
    2. 山东省海空信息感知与处理技术重点实验室,山东 烟台 264001
    3. 哈尔滨工程大学信息与通信工程学院,黑龙江 哈尔滨 150001
  • 收稿日期:2025-07-14 修回日期:2025-09-25 出版日期:2026-01-24 发布日期:2026-01-24
  • 通讯作者: 毛忠阳 E-mail:lxg1023@163.com
  • 作者简介:刘锡国(1981—),男,副教授,博士,主要研究方向为智能通信抗干扰、无线光通信
    刘 敏(1983—),女,教授,博士,主要研究方向为水声通信与信号处理、智能决策
    徐盟飞(1999—),男,硕士研究生,主要研究方向为信息感知与处理
    马 鹏(2000—),男,硕士研究生,主要研究方向为通信抗干扰决策
  • 基金资助:
    国家自然科学基金(U23A20271)资助课题

Anti-jamming decision-making algorithm based on information-regulated LSTM

Xiguo Liu1,2(), Min Liu1,2, Zhongyang Mao1,2, Mengfei Xu3, Peng Ma3   

  1. 1. Naval Aviation university,Yantai 264001, China
    2. Key Laboratory of Sea and Air Information Sensing and Processing Technology of Shandong Porvincial,Yantai 264001,China
    3. College of Information and Communication Engineering,Harbin Engineering University,Harbin 150001,China
  • Received:2025-07-14 Revised:2025-09-25 Online:2026-01-24 Published:2026-01-24
  • Contact: Zhongyang Mao E-mail:lxg1023@163.com

摘要:

经典强化学习算法在抗干扰决策领域虽然具有一定的作用,但收敛速度较慢,难以在短时间内生成有效的抗干扰策略,这在实际应用中是难以接受的。为提高抗干扰能力的同时尽量加快其收敛速度,基于强化学习算法框架,使用信息调节长短期记忆网络构建策略和价值网络,高效建模干扰环境中的频谱状态,筛选重要特征、抑制无效信息,提高了模型对环境的建模准确度。同时,使用经验补充机制加速对算法经验池的填充,促使算法快速达到训练状态。两者结合,使得算法的归一化吞吐量到达0.65时,使用的回合数比柔性演员-评论员(soft actor-critic,SAC)算法少85.3%,算法收敛后的归一化吞吐量比SAC算法高出26.3%。实验结果表明,所提算法对环境的建模准确度高、收敛速度快、归一化吞吐量更高,表现出良好的适应性和鲁棒性。

关键词: 抗干扰决策, 强化学习, 深度强化学习, 收敛速度, 经验补充

Abstract:

Although classical reinforcement learning algorithms have shown certain effectiveness in the field of anti-jamming decision-making, their convergence speed is typically slow, and it is difficult to generate effective anti-jamming strategies in a short period of time, which is unacceptable in practical applications. To enhance anti-jamming capability while accelerating convergence, a reinforcement learning framework that incorporates an information-regulated long short-term memory is proposed to construct both the policy and value networks. This architecture efficiently models the spectral state under jamming conditions by filtering salient features and suppressing irrelevant information, thereby improving the model’s environmental representation accuracy. Additionally, an experience augmentation mechanism is introduced to accelerate the population of the replay buffer, enabling the algorithm to reach the training-ready state more quickly. Experimental results demonstrate that when the normalized throughput reaches 0.65, the proposed method requires 85.3% fewer episodes than the standard soft actor-critic (SAC) algorithm. Moreover, the normalized throughput after convergence exceeds that of SAC by 26.3%. Experimental results show that the proposed algorithm has high modeling accuracy, high convergence rate, fast normalized throughput and better adaptability and robustness.

Key words: anti-jamming decision-making, reinforcement learning, deep reinforcement learning, convergence speed, experience augmentation

中图分类号: