To solve the problems faced by the existing edge computing task scheduling based on deep reinforcement learning, such as fixed action space exploration, low sample efficiency, large memory demand and poor stability and to better carry out effective task scheduling in the edge computing system with relatively limited computing resources, an adaptive edge computing task scheduling method D3DQN-CAA is proposed based on the improved deep reinforcement learning model D3DQN (Dueling Double DQN). In the task offloading decision, the corresponding relationship between the task and processor is regarded as a multidimensional knapsack problem, and the computing node with the highest matching degree is selected for task processing according to the state information of the current scheduled task and the computing node; For improving the parameters updating efficiency of the evaluation network and reducing the influence of overestimation, a comprehensive Q-value calculation method is proposed; For accelerating the convergence speed of neural networks, an adaptive dynamic exploration degree of action space adjustment strategy is proposed; For reducing the storage resources required and improving the sample efficiency, an adaptive lightweight prioritized playback mechanism is proposed. Experimental results show that compared with multiple benchmark algorithms, the D3DQN-CAA algorithm can effectively reduce the number of training steps of deep reinforcement learning networks and make full use of edge computing resources to improve the real-time performance of task processing and reduce the system energy consumption.
系统整体处理流程如图8所示.处理开始时,初始化模型各项参数,itr为1;若迭代次数满足条件1:itr mod N1=0为参数复制周期(N1),则复制参数到目标网络.若迭代次数满足条件2:itr mod N2=0为经验回放周期(N2),则回放经验池中的学习经验并清空经验池.不满足以上条件,则获取环境状态信息.获得评估网络和目标网络的输出后,利用综合性Q值计算方法CQC获得最终输出;并通过自适应动作空间动态探索度调整策略AGP生成卸载决策;根据当前任务的卸载决策将任务卸载至对应的计算节点进行处理或本地处理,并计算相应的加权开销;再计算损失值,启动自适应轻量优先级回放机制ALPR中的存储经验功能,存储优先级最高的学习经验;将损失值放入存储序列中,若满足条件3:itr mod N3=0,为损失值均值替换周期(N3),则计算存储序列中的损失值均值,替换AGP中均值并清空序列;基于损失值更新评估网络参数.
ZHOUY Z, ZHANGD .Near-end cloud computing:opportunities and challenges in the post-cloud computing era[J].Chinese Journal of Computers,2019,42(4):677-700.(in Chinese)
[3]
WANGF X, ZHANGM, WANGX X,et al .Deep learning for edge computing applications:a state-of-the-art survey[J].IEEE Access,2020,8:58322-58336.
[4]
SHIW S, CAOJ, ZHANGQ,et al .Edge computing:vision and challenges[J].IEEE Internet of Things Journal,2016,3(5):637-646.
[5]
WANGX F, HANY W, LEUNGV C M,et al .Convergence of edge computing and deep learning:a comprehensive survey[J].IEEE Communications Surveys & Tutorials,2020,22(2):869-904.
LUH F, GUC H, LUOF,et al .Research on task offloading based on deep reinforcement learning in mobile edge computing[J].Journal of Computer Research and Development,2020,57(7):1539-1554.(in Chinese)
ZHANGZ Y, CHENY F, WANGY H,et al .Multi-robot task allocation algorithm b Multirobot task allocation algorithm based on heuristically accelerated deep Q network[J].Journal of Harbin Engineering University,2022,43(6):857-864.(in Chinese)
YUP, ZHANGJ Y, LIW J,et al. Energy-efficient resource allocation method in mobile edge network based on double deep Q-learning[J].Journal on Communications,2020,41(12):148-161.(in Chinese)
[12]
ZHUA Q, GUOS T, MAM F,et al .Computation offloading for workflow in mobile edge computing based on deep Q-learning[C]//2019 28th Wireless and Optical Communications Conference (WOCC).Beijing,China: IEEE,2019:1-5.
[13]
TANGM, WONGV W S .Deep reinforcement learning for task offloading in mobile edge computing systems[J].IEEE Transactions on Mobile Computing,2022,21(6):1985-1997.
[14]
HANB A, YANGJ J .Research on adaptive job shop scheduling problems based on dueling double DQN[J].IEEE Access,2020,8:186474-186495.
[15]
XIONGX, ZHENGK, LEIL,et al .Resource allocation based on deep reinforcement learning in IoT edge computing[J].IEEE Journal on Selected Areas in Communications,2020,38(6):1133-1146.
[16]
ZOUJ F, HAOT B, YUC,et al .A3C-DO:a regional resource scheduling framework based on deep reinforcement learning in edge scenario[J].IEEE Transactions on Computers,2021, 70(2): 228-239.
[17]
QIF, ZHUOL, XINC .Deep reinforcement learning based task scheduling in edge computing networks[C]//2020 IEEE/CIC International Conference on Communications in China (ICCC).Chongqing,China: IEEE,2020:835-840.
[18]
NATHS, WUJ X .Deep reinforcement learning for dynamic computation offloading and resource allocation in cache-assisted mobile edge computing systems[J].Intelligent and Converged Networks,2020,1(2):181-198.
[19]
KEH C, WANGJ, DENGL Y,et al .Deep reinforcement learning-based adaptive computation offloading for MEC in heterogeneous vehicular networks[J].IEEE Transactions on Vehicular Technology,2020,69(7):7916-7929.
[20]
NATHS, WUJ X. Dynamic computation offloading and resource allocation for multi-user mobile edge computing[C]//GLOBECOM 2020—2020 IEEE Global Communications Conference.Taipei,China: IEEE,2020:1-6.
JIANGW L, WUJ, WANGY N .Autonomous obstacle avoidance and target tracking of UAV based on meta-reinforcement learning[J].Journal of Hunan University (Natural Sciences),2022, 49(6):101-109.(in Chinese)
CHENZ, JIANGW H, DUJ W .Research on deep reinforcement learning sequential recommendation algorithm based on policy memory[J].Journal of Hunan University (Natural Sciences),2022,49(8):208-216.(in Chinese)
[25]
VAN HASSELTH, GUEZA, SILVERD .Deep reinforcement learning with double Q-learning[J].Proceedings of the AAAI Conference on Artificial Intelligence,2016,30(1):129-144.
[26]
MNIHV, KAVUKCUOGLUK, SILVERD,et al .Human-level control through deep reinforcement learning[J].Nature,2015,518(7540):529-533.
[27]
LIF C, HUB. DeepJS:job scheduling based on deep reinforcement learning in cloud data center[C]//Proceedings of the 2019 4th International Conference on Big Data and Computing -ICBDC 2019.May 10-12,2019.Guangzhou,China: ACM,2019:48-53.
[28]
ARABNEJADH, BARBOSAJ G .List scheduling algorithm for heterogeneous systems by an optimistic cost table[J].IEEE Transactions on Parallel and Distributed Systems,2014,25(3):682-694.