针对目前矿井传感器所收集数据的传输效率差、实时性低、丢包率高等问题,提出了一种基于深度强化学习的无人机矿井自主巡航解决方法,以有效收集物联网节点数据。该方法以无人机作为传输中介,根据矿井物联网节点数据生成周期性的不同差值,利用强化学习TD3(Twin Delayed Deep Deterministic policy gradient algorithm)算法实现无人机最优路径规划。同时算法考虑并设计了符合矿井实际场景的环境、奖励值、状态信息等。提出了一种预测等待的方法,预测待采集数据产生时间并确定目标节点,无人机在信号覆盖范围内前往目标节点提前等待,以实时获取矿井传感器的生成数据。实验结果表明,无人机能够自主决策实现最优路径规划,并收集节点数据;在训练回合为700时,奖励值达到峰值,算法达到收敛并具备优异的表现。
Abstract
Aiming at the problems of poor transmission efficiency, low real-time performance and high packet loss rate of data collected by mine sensors at present, a mine autonomous cruise method of unmanned aerial vehicle(UAV) based on deep reinforcement learning was proposed to effectively collect data of Internet of Things nodes. Specifically, the method takes UAV as the transmission intermediary, and according to the different values of data generation cycle of mine IOT node, Twin Delayed Deep Deterministic policy gradient algorithm (TD3) is used to realize the optimal path planning of UAV. At the same time, the algorithm also considers and designs the environment, reward value, and state information that conforms to the actual mine scene. In particular, we propose a predictive waiting method, which predicts the generation time of data to be collected and determines the target node. The UAV goes to the target node and waits in advance within the signal coverage range to obtain the data generated by the mine sensor in real time. The experimental results show that the UAV can realize the optimal path planning and collect node data by autonomous decision. When the training round is 700, the reward value reaches the peak, the algorithm reaches convergence and has excellent performance.
现有的基于物联网数据采集的无人机路径规划方案中,大多采用LOS通道,即视距传播,虽有助于综合系统测试及通信资源优化,但通信质量及吞吐量受障碍物的影响,不利于复现复杂环境动态模拟。概率LOS模型的提出[17],从概率的角度综合考虑了LOS与非视距(non line of sight,NLOS),但实际环境中,无人机与传感器节点之间并非没有阻碍,与实际需求不符。Zeng等[18]提出了考虑实际环境的城市LOS模型,但只考虑了单独的传感器,不适用于多传感器分布的情况。
KIMM L, PEVZNERL D, TEMKINI O. Development of automatic system for Unmanned Aerial Vehicle (UAV) motion control for mine conditions[J]. Mining Science and Technology (Russia), 2021, 6(3): 203-210. DOI: 10.17073/2500-0632-2021-3-203-210 .
[2]
HEßELINGC, KELLERJ. Pareto-optimal covert channels in sensor data transmission[C]//Proceedings of the 2022 European Interdisciplinary Cybersecurity Conference. New York: ACM, 2022: 79-84. DOI: 10.1145/3528580.3532844 .
[3]
LIUT, NIS G, LIX Q, et al. Deep reinforcement learning based approach for online service placement and computation resource allocation in edge computing[J]. IEEE Transactions on Mobile Computing,2022,PP(99):1.DOI: 10.1109/TMC.2022.3148254 .
ZHANGD, WUP L, ZHENGX Z, et al. Research status and development trend of mine detection unmanned aerial vehicle[J].Industry and Mine Automation, 2020,46(7): 76-81.DOI: 10.13272/j.issn.1671-251x.17538 (Ch ).
[6]
ZENGY, ZHANGR, LIMT J. Wireless communications with unmanned aerial vehicles: Opportunities and challenges[J]. IEEE Communications Magazine, 2016, 54(5):36-42. DOI: 10.1109/MCOM.2016.7470933 .
[7]
WUC X, JUB B, WUY, et al. UAV autonomous target search based on deep reinforcement learning in complex disaster scene[J]. IEEE Access, 2019, 7: 117227-117245. DOI: 10.1109/ACCESS.2019.2933002 .
[8]
YANC, XIANGX J, WANGC. Towards real-time path planning through deep reinforcement learning for a UAV in dynamic environments[J]. Journal of Intelligent & Robotic Systems, 2020, 98(2): 297-309. DOI: 10.1007/s10846-019-01073-3 .
[9]
WANGY, GAOZ, ZHANGJ, et al. Trajectory design for UAV-based Internet of Things data collection: A deep reinforcement learning approach[J]. IEEE Internet of Things Journal, 2022, 9(5): 3899-3912. DOI: 10.1109/JIOT.2021.3102185 .
[10]
LIS D, XUX, ZUOL. Dynamic path planning of a mobile robot with improved Q-learning algorithm[C]//2015 IEEE International Conference on Information and Automation. New York: IEEE Press, 2015: 409-414. DOI: 10.1109/ICInfA.2015.7279322 .
[11]
SHIK J, WUP, LIUM S. Research on path planning method of forging handling robot based on combined strategy[C]//2021 IEEE International Conference on Power Electronics,Computer Applications (ICPECA). New York: IEEE Press, 2021: 292-295. DOI: 10.1109/ICPECA51329.2021.9362595 .
[12]
KHANA I, AL-MULLAY. Unmanned aerial vehicle in the machine learning environment[J]. Procedia Computer Science, 2019, 160(C): 46-53. DOI: 10.1016/j.procs.2019.09.442 .
[13]
SINGLAA, PADAKANDLAS, BHATNAGARS. Memory-based deep reinforcement learning for obstacle avoidance in UAV with limited environment knowledge[J]. IEEE Transactions on Intelligent Transportation Systems, 2021, 22(1): 107-118. DOI: 10.1109/TITS.2019.2954952 .
[14]
WANGD W, FANT X, HANT, et al. A two-stage reinforcement learning approach for multi-UAV collision avoidance under imperfect sensing[J]. IEEE Robotics and Automation Letters, 2020, 5(2): 3098-3105. DOI: 10.1109/LRA.2020.2974648 .
[15]
DINGR J, GAOF F, SHENX S. 3D UAV trajectory design and frequency band allocation for energy-efficient and fair communication: A deep reinforcement learning approach[J]. IEEE Transactions on Wireless Communications, 2020, 19(12): 7796-7809. DOI: 10.1109/TWC.2020.3016024 .
[16]
LIY L, FANGH W, LIM Y, et al. Neural network pruning and fast training for DRL-based UAV trajectory planning[C]//2022 27th Asia and South Pacific Design Automation Conference (ASP-DAC).New York: IEEE Press, 2022: 574-579. DOI: 10.1109/ASP-DAC52403.2022.9712561 .
[17]
ZHAOJ J, YUL, CAIK Q, et al. RIS-aided ground-aerial NOMA communications: A distributionally robust DRL approach[J]. IEEE Journal on Selected Areas in Communications, 2022, 40(4): 1287-1301. DOI: 10.1109/JSAC.2022.3143230 .
[18]
YANGP, CAOX B, XIX, et al. Three-dimensional continuous movement control of drone cells for energy-efficient communication coverage[J]. IEEE Transactions on Vehicular Technology, 2019, 68(7): 6535-6546. DOI: 10.1109/TVT.2019.2913988 .
[19]
ZENGY, XUX L. Path design for cellular-connected UAV with reinforcement learning[C]//2019 IEEE Global Communications Conference (GLOBECOM). New York: IEEE Press, 2020: 1-6. DOI: 10.1109/GLOBECOM38437.2019.9014041 .