| 1 |
孙宇祥, 彭益辉, 李斌, 等. 智能博弈综述: 游戏AI对作战推演的启示[J]. 智能科学与技术学报, 2022, 4 (2): 157.
|
| 2 |
Zhang N, Zhu X, Gou Y. The role of artificial intelligence and autonomous systems in the decision-making center of the US mosaic war[C]// International Conference on Intelligent Computing and Human-Computer Interaction, 2020: 21.
|
| 3 |
况立群, 李思远, 冯利, 等. 深度强化学习算法在智能军事决策中的应用[J]. 计算机工程与应用, 2021, 57 (20): 271.
|
| 4 |
蒲志强, 易建强, 刘振, 等. 知识和数据协同驱动的群体智能决策方法研究综述[J]. 自动化学报, 2022, 48 (3): 627.
|
| 5 |
袁博文, 刘东波, 刘兆鹏, 等. 基于改进行为树的作战计划建模方法[J]. 系统工程与电子技术, 2023, 45 (4): 1111.
|
| 6 |
Lecun Y, Bengio Y, Hinton G. Deep learning[J]. Nature, 2015, 521 (7553): 436.
|
| 7 |
Mnih V, Kavukcuoglu K, Silver D, et al. Human-level control through deep reinforcement learning[J]. Nature, 2015, 518 (7540): 529.
|
| 8 |
刘全, 翟建伟, 章宗长, 等. 深度强化学习综述[J]. 计算机学报, 2018, 41 (1): 1.
|
| 9 |
孙长银, 穆朝絮. 多智能体深度强化学习的若干关键科学问题[J]. 自动化学报, 2020, 46 (7): 1301.
|
| 10 |
Schrittwieser J, Antonoglou I, Hubert T, et al. Mastering Atari, Go, chess and shogi by planning with a learned model[J]. Nature, 2020, 588 (7839): 604.
|
| 11 |
Hernandez-leal P, Kartal B, Taylor M E. A survey and critique of multiagent deep reinforcement learning[J]. Autonomous Agents and Multi-Agent Systems, 2019, 33 (6): 750.
|
| 12 |
梁星星, 冯旸赫, 马扬, 等. 多Agent深度强化学习综述[J]. 自动化学报, 2020, 46 (12): 2537.
|
| 13 |
Vinyals O, Babuschkin I, Czarnecki W M, et al. Grandmaster level in StarCraft II using multi-agent reinforcement learning[J]. Nature, 2019, 575 (7782): 350.
|
| 14 |
Boron J, Darken C. Developing combat behavior through reinforcement learning in wargames and simulations[C]// IEEE Conference on Games, 2020: 728.
|
| 15 |
Cannon C T, Goericke S. Using convolution neural networks to develop robust combat behaviors through reinforcement learning[D]. Monterey: Naval Postgraduate School, 2021.
|
| 16 |
Tarraf D C, Gilmore J M, Barnett D S, et al. An experiment in tactical wargaming with platforms enabled by artificial intelligence[J]. The Journal of Defense Modeling and Simulation, 2022, 19 (4): 369.
|
| 17 |
李理, 李旭光, 郭凯杰, 等. 国产化环境下基于强化学习的地空协同作战仿真[J]. 兵工学报, 2022, 43 (S1): 74.
|
| 18 |
施伟, 冯旸赫, 程光权, 等. 基于深度强化学习的多机协同空战方法研究[J]. 自动化学报, 2021, 47 (7): 1610.
|
| 19 |
郭洪宇, 初阳, 刘志, 等. 基于深度强化学习潜艇攻防对抗训练指挥决策研究[J]. 指挥控制与仿真, 2022, 44 (1): 103.
|
| 20 |
徐志雄, 曹雷, 陈希亮. 基于元深度强化学习方法的智能博弈决策模型研究[J]. 军事运筹与系统工程, 2021, 35 (3): 66.
|
| 21 |
Kong W R, Zhou D Y, Zhao Y Y, et al. Maneuvering strategy generation algorithm for multi-UAV in close-range air combat based on deep reinforcement learning and self-play[J]. Control Theory & Applications, 2022, 39 (2): 352.
|
| 22 |
Sutton R S, Barto A G. Reinforcement learning: an introduction[M]. 2nd ed. Cambridge: MIT Press, 2018.
|
| 23 |
Silver D, Singh S, Precup D, et al. Reward is enough[J]. Artificial Intelligence, 2021, 299, 103535.
|
| 24 |
Wang X J, Song J X, Qi P H, et al. SCC: an efficient deep reinforcement learning agent mastering the game of StarCraft II[C]// International Conference on Machine Learning, 2021: 10905.
|
| 25 |
Vaswani A, Shazeer N, Parmar N, et al. Attention is all you need[J]. Advances in Neural Information Processing Systems, 2017, 30:6000.
|
| 26 |
Kim J, El-Khamy M, Lee J. Residual LSTM: design of a deep recurrent architecture for distant speech recognition[EB/OL]. [2024-08-08]. https://arxiv.org/abs/1701.03360.
|
| 27 |
Vinyals O, Fortunato M, Jaitly N. Pointer networks[J]. Advances in Neural Information Processing Systems, 2015, 28, 2692.
|
| 28 |
Yu C, Velu A, Vinitsky E, et al. The surprising effectiveness of PPO in cooperative multi-agent games[J]. Advances in Neural Information Processing Systems, 2022, 35, 24611.
|
| 29 |
Espeholt L, Soyer H, Munos R, et al. Impala: scalable distributed deep-RL with importance weighted actor-learner architectures[C]// International Conference on Machine Learning, 2018: 1407.
|
| 30 |
Zhao L Y, Chang T Q, Zhang J, et al. A policy optimization algorithm based on sample adaptive reuse and dual-clipping for robotic action control[J]. Applied Soft Computing, 2023, 134, 109967.
|
| 31 |
Schulman J, Wolski F, Dhariwal P, et al. Proximal policy optimization algorithms[EB/OL]. [2024-08-08]. https://arxiv.org/abs/1707.06347.
|
| 32 |
Haarnoja T, Zhou A, Abbeel P, et al. Soft actor-critic: off-policy maximum entropy deep reinforcement learning[C]// 35th International Conference on Machine Learning, 2018: 1861.
|