1.College of Electrical and Information Engineering,Hunan University,Changsha 410082,China
2.Greater Bay Area Institute for Innovation,Hunan University,Guangzhou 511300,China
Show less
文章历史+
Received
Published
2024-11-28
2026-04-25
Issue Date
2026-06-11
PDF (2271K)
摘要
针对深度强化学习算法在求解复杂环境下动态目标的无人机智能追踪问题时存在的收敛速度慢、成功率低、模型泛用性差等问题, 将注意力机制与深度确定性策略梯度(deep deterministic policy gradient, DDPG)结合, 提出了attention-DDPG模型, 并在此基础上结合多经验池方法, 搭建了一种新的深度强化学习算法——多经验池注意力深度确定性策略梯度(multi pool attention deep deterministic policy gradient, MPADDPG)算法. 算法将注意力机制加入 DDPG 的 Actor 网络, 赋予状态中各个分量不同的权重来突出重要且关键的信息, 并引入多经验池机制分离失败经验, 强化成功经验, 提升了算法的收敛性;更进一步, 通过赋予无人机环境感知能力, 提升了算法的泛化能力. 最后, 在建立的无人机的连续状态空间与动作空间中验证了MPADDPG算法的有效性.仿真结果表明,MPADDPG算法的智能追踪成功率超过90%, 较DDPG算法具有更高的追踪成功率与更强的泛化能力.
Abstract
To address the challenges of slow convergence rates, low success rates, and poor model generalization in deep reinforcement learning algorithms for UAV intelligent tracking decision-making under a complex environment, this study combines the attention mechanism with the deep deterministic policy gradient (DDPG) approach and introduces an attention-DDPG model. Meanwhile, incorporating a multi-experience pool method, a proposed algorithm, named multi pool attention deep deterministic policy gradient (MPADDPG) is built up, and the attention mechanism is integrated into DDPG’s Actor network. This allows for different weights to be assigned to various state components, so as to highlight the crucial information, while a multi-experience pool strategy is employed to distinguish unsuccessful experiences from successful ones in order to improve convergence. Additionally, enhancing the UAV’s environmental perception capabilities further boosts the algorithm’s generalization ability. The effectiveness of MPADDPG is validated within a continuous state and action space framework established in this research. Simulation results demonstrate that MPADDPG achieves an intelligent tracking success rate exceeding 90%, outperforming DDPG in both tracking success and generalization capability.
JIANGW L, WUJ, WANGY N .Autonomous obstacle avoidance and target tracking of UAV based on meta-reinforcement learning[J].Journal of Hunan University (Natural Sciences), 2022, 49(6):101-109.(in Chinese)
FANR T, LIUH, CHENGM, et al. Game environment of multiple unmanned aerial vehicle system coordinated path planning [J/OL]. Journal of Beijing University of Aeronautics and Astronautics,1-10[2024-11-28].in Chinese)
JIANGK, CAOJ Y, LIUW Z, et al .Research on reinforcement learning methods for navigation and adversarial control in mobile robots[J].Control Theory & Applications, 2025, 42(9):1757-1765.(in Chinese)
LOUZ B, PENGY, XINK .Research on path planning system based on improved Q-Learning algorithm[J]. Computer and Digital Engineering, 2024, 52(8): 2312-2316.(in Chinese)
WANGJ, WANGG P, SUNT Y. Automatic planning model of UAV campus security monitoring path based on DQN algorithm[J]. Automation & Instrumentation, 2024(4):193-196.(in Chinese)
ZHAOQ, ZHENZ Y, GONGH J, et al .UAV formation control based on dueling double DQN[J].Journal of Beijing University of Aeronautics and Astronautics, 2023, 49(8):2137-2146.(in Chinese)
BIW H, DUANX B. Research on UAV path planning method based on deep reinforcement learning[J]. Aeronautical Science & Technology, 2023, 34(12): 118-124.(in Chinese)
WANGZ Y, LIUZ, LIY W, et al .Research on replenishment of urban UAV distribution based on artificial potential field method and binocular vision under the background of Internet of Things[J].Information Recording Materials, 2024, 25(3):240-242.(in Chinese)
ZHANW C, GUOL J, XUS J, et al .UAV Intelligent obstacle avoidance algorithm based on greedy DDPG[J].Journal of Air & Space Early Warning Research, 2024, 38(5):342-346.(in Chinese)
RONGC T, ZHUH W, ZHANGB, et al .Research on path planning of mobile robots based on deep reinforcement learning[J].Modern Informationn Technology, 2024, 8(16):60-63.(in Chinese)
CHENZ, JIANGW H, DUJ W .Research on deep reinforcement learning sequential recommendation algorithm based on policy memory[J].Journal of Hunan University (Natural Sciences), 2022, 49(8):208-216.(in Chinese)