To address the considerable challenges posed by the randomness and uncertainty of customer demand for fresh produce-demand that fluctuates with personal preferences, emergencies, and other external factors-this study integrates reinforcement learning with robust optimisation to build an intelligent decision-making system capable of operating effectively under demand uncertainty. First, a Copula-Gamma model is employed to generate highly correlated demand scenarios, thereby capturing inter-customer dependence more precisely and supplying realistic, dynamic inputs for route planning. These scenarios are then fed into a Dueling-Double DQN to obtain an initial set of delivery paths, after which second-order cone programming (SOCP)-based robust optimisation is introduced to regulate excess costs, mitigating the cost volatility induced by demand uncertainty. The results showed that: 1) Under demand uncertainty, the proposed Dueling-Double DQN algorithm significantly reduces the average total cost across all instances in the R, C and RC groups of the Solomon dataset, compared to the baseline generated by the heuristic nearest-neighbour greedy algorithm, as well as Vanilla DQN and PPO. 2) With robust optimization, the second-stage random cost drops sharply-by 72.09% in Group R, 70.80% in Group RC, and 67.02% in Group C. In summary, the proposed method not only effectively solves the cost optimisation problem for fresh product delivery under demand uncertainty, but also enhances delivery stability, providing a generalisable and intelligent optimisation solution for practical delivery systems.
针对生鲜产品无人机配送场景,构建一个考虑不确定需求与运输损耗的生鲜农产品无人机配送路径优化模型。研究的核心问题可描述为:如何为一组无人机车队规划其访问所有客户点的行驶路径,面对需求不确定时,在满足无人机运力、路径连通性等物理约束的前提下,协同优化无人机固定成本、运输成本和生鲜货损成本,从而实现系统总成本的最小化。该问题本质上是需求不确定的车辆路径问题(Vehicle routing problem with stochastic demands,VRPSD)。
为进一步分析不同参数α对鲁棒优化模型的影响,本研究采用变异系数(Coefficient of variation,CV)方法分析了不同场景情况下α对总成本的影响,结果如图1所示。可知:当α由0.1增加至0.5时,CV增长缓慢,系统鲁棒性逐步增强,且总成本上升幅度较小,表明模型在提高稳定性的同时仍保持较好的成本效率。然而,当α>0.5时,CV和总成本均呈现陡增趋势,说明鲁棒性提升已开始带来明显的效率损失,即模型变得过于保守。故α=0.5附近表现出较好的综合平衡,验证了所提模型在不同客户分布场景下的适应性。
MaC X, XueF S, MaC R, LiH J. Route optimization of fresh food distribution under time-varying network and hybrid adjustment strategy[J]. Journal of Transportation Systems Engineering and Information Technology, 2023, 23(4): 298-306 (in Chinese)
ZhangJ F, YangZ H. Research on distribution path optimization of multi-temperature cold chain in time-varying road network environment [J]. Journal of Chongqing Normal University: Natural Science, 2020, 37(1): 119-126 (in Chinese)
CaiW G, LiuJ X, ZhangX X. Algorithm for taxi ride-sharing scheduling based on probabilistic routing[J]. Application Research of Computers, 2024, 41(2): 432-437 (in Chinese)
LiY, FanH M, ZhangX N, YangX. Two-phase variable neighborhood scatter search for the capacitated vehicle routing problem with stochastic demand[J]. Control Theory & Applications, 2017, 34(12): 1594-1604 (in Chinese)
[15]
ReuskenM, LaporteG, RohmerS U K, CruijssenF. Vehicle routing with stochastic demand, service and waiting times: The case of food bank collection problems [J]. European Journal of Operational Research, 2024, 317(1): 111-127
[16]
CaiH, XuP, TangX, LinG. Solving the vehicle routing problem with stochastic travel cost using deep reinforcement learning[J]. Electronics, 2024, 13(16): 3242
[17]
ZhouC H, MaJ X, DougeL, ChewE P, LeeL H. Reinforcement Learning-based approach for dynamic vehicle routing problem with stochastic demand[J]. Computers & Industrial Engineering, 2023, 182: 109443
[18]
ShahryariE, ShayeghiH, Mohammadi-ivatlooB, MoradzadehM. A copula-based method to consider uncertainties for multi-objective energy management of microgrid in presence of demand response [J]. Energy, 2019, 175: 879-890
[19]
WangZ Y, SchaulT, HesselM, Van HasseltH, LanctotM, De FreitasN. Dueling network architectures for deep reinforcement learning[C]. In: Proceedings of the 33rd International Conference on Machine Learning. London: Google DeepMind, 2016(48): 1995-2003
[20]
LiJ W, MaY N, GaoR Z, CaoZ G, LimA, SongW, ZhangJ. Deep reinforcement learning for solving the heterogeneous capacitated vehicle routing problem[J]. IEEE Transactions on Cybernetics, 2022, 52(12): 13572-13585
[21]
Van HasseltH, GuezA, SilverD. Deep reinforcement learning with double Q-Learning[C]. In: Proceedings of the Thirtieth AAAI Conference on Artificial Intelligence. London: Google DeepMind, 2016: 2094-2100
[22]
RuiG, AntonK. Distributionally robust stochastic optimization with wasserstein distance[J]. Mathematics of Operations Research, 2022, 48(2): 603-655
[23]
EsfahaniM P, KuhnD. Data-driven distributionally robust optimization using the Wasserstein metric: performance guarantees and tractable reformulations [J]. Mathematical Programming, 2018, 171(1/2): 115-166