基于GAIL-PPO两阶段训练的高速公路车辆换道决策优化框架

郭洪江 ,  朱政泽 ,  龚家元 ,  殷政 ,  吕成志

工业工程 ›› 2026, Vol. 29 ›› Issue (4) : 96 -105.

PDF (1137KB)
工业工程 ›› 2026, Vol. 29 ›› Issue (4) : 96 -105. DOI: 10.3969/j.issn.1007-7375.250150
系统建模与优化

基于GAIL-PPO两阶段训练的高速公路车辆换道决策优化框架

作者信息 +

A Highway Vehicle Lane-Changing Decision Optimization Framework Based on Two-stage GAIL-PPO Training

Author information +
文章历史 +
PDF (1163K)

摘要

针对高速公路场景下智能车辆换道决策中模仿学习对数据质量依赖性高、泛化能力有限以及强化学习训练效率低、多目标难以兼顾等问题,提出一种基于生成对抗模仿学习 (generative adversarial imitation learning, GAIL) 与近端策略优化 (proximal policy optimization, PPO) 的两阶段协同优化框架。在 GAIL 判别器中引入基于 Wasserstein 距离并结合梯度惩罚的对抗机制以提升训练稳定性;并将 PPO 纳入 GAIL 的生成器更新过程,通过 Actor-Critic 架构增强策略学习的鲁棒性;采用 PPO 对预训练策略进行多目标强化微调,构建融合通行效率与安全约束的多目标奖励函数,实现从专家模仿到复杂场景约束下的策略优化。基于 highway-env 仿真环境的实验结果表明,与 PPO baseline 和 DQN 方法相比,所提出方法在平均通行速度方面分别提升约 4% 和 8%,同时有效减少不必要的换道行为;结合纵向加速度时间序列分析与鲁棒性测试结果,进一步验证了该方法在不同驾驶时长、交通流密度、车道数及车辆动力学约束条件下的稳定性与泛化能力。

Abstract

This paper addresses challenges in intelligent vehicle lane-changing decision-making on highways, including the high dependency of imitation learning on data quality and its limited generalization capability, as well as the low training efficiency of reinforcement learning and its difficulty in balancing multiple objectives. To this end, we propose a two-stage collaborative optimization framework based on Generative Adversarial Imitation Learning (GAIL) and Proximal Policy Optimization (PPO). A Wasserstein distance-based adversarial mechanism with gradient penalty is introduced into the GAIL discriminator to enhance training stability. PPO is integrated into the generator update process of GAIL through an actor–critic architecture, improving the robustness of policy learning. The pre-trained policy is then fine-tuned using PPO-based multi-objective reinforcement learning with a multi-objective reward function that balances traffic efficiency and safety constraints. This enables the transition from expert imitation to policy optimization under complex scenario constraints. Experimental results in the highway-env simulation environment demonstrate that the proposed approach improves average travel speed by approximately 4% and 8% compared with the PPO baseline and DQN, respectively, while effectively reducing unnecessary lane changes. Further analysis of longitudinal acceleration time series and robustness tests validates the method's stability and generalization capabilities under different driving durations, traffic densities, lane numbers, and vehicle dynamics.

关键词

高速换道决策 / 生成对抗模仿学习 / 深度强化学习 / 近端策略优化

Key words

highway lane-changing decision-making / generative adversarial imitation learning(GAIL) / deep reinforcement learning / proximal policy optimization(PPO)

引用本文

引用格式 ▾
郭洪江,朱政泽,龚家元,殷政,吕成志. 基于GAIL-PPO两阶段训练的高速公路车辆换道决策优化框架[J]. 工业工程, 2026, 29(4): 96-105 DOI:10.3969/j.issn.1007-7375.250150

登录浏览全文

4963

注册一个新账户 忘记密码

参考文献

[1]

崔冰艳, 李贺, 崔哲, . 智能网联汽车换道决策安全性研究综述[J]. 交通信息与安全, 2023, 41(4): 1-13.

[2]

Cui Bingyan, Li He, Cui Zhe, et al. A review of safety studies on lane change decision-makings for connected automated vehicles[J]. Journal of Transport Information and Safety, 2023, 41(4): 1-13.

[3]

Mechernene A, Judalet V, Chaibet A, et al. Lane change decision algorithm based on risk prediction and fuzzy logic method[C]// 2021 25th International Conference on System Theory, Control and Computing (ICSTCC). Piscataway: IEEE, 2021: 707-713.

[4]

Toledo T, Koutsopoulos H N, Ben-Akiva M. Integrated driving behavior modeling[J]. Transportation Research Part C: Emerging Technologies, 2007, 15(2): 96-112.

[5]

张可琨, 曲大义, 宋慧, . 自动驾驶车辆换道博弈策略分析及建模[J]. 复杂系统与复杂性科学, 2023, 20(2): 60-67.

[6]

Zhang Kekun, Qu Dayi, Song Hui, et al. Analysis and modeling for lane-changing game strategy of autonomous vehicles[J]. Complex Systems and Complexity Science, 2023, 20(2): 60-67.

[7]

Li W B, Wang W D, Yang C, et al. A mandatory lane changing integrated decision making and planning method using game theory for autonomous vehicle[C]// 2023 7th CAA International Conference on Vehicular Control and Intelligence (CVCI). Piscataway: IEEE, 2024: 1-8.

[8]

罗鹏, 黄珍, 秦易晋, . 基于 DQN 的车辆驾驶行为决策方法[J]. 交通信息与安全, 2020, 38(5): 67-77,112.

[9]

Luo Peng, Huang Zhen, Qin Yijin, et al. A method of vehicle driving behavior decision based on DQN algorithm[J]. Journal of Transport Information and Safety, 2020, 38(5): 67-77,112.

[10]

宋晓琳, 盛鑫, 曹昊天, . 基于模仿学习和强化学习的智能车辆换道行为决策[J]. 汽车工程, 2021, 43(1): 59-67.

[11]

Song Xiaolin, Sheng Xin, Cao Haotian, et al. Lane-change behavior decision-making of intelligent vehicle based on imitation learning and reinforcement learning[J]. Automotive Engineering, 2021, 43(1): 59-67.

[12]

Guo L, Liu X Z. Lane-changing decisions making for autonomous vehicles via behavior cloning and decision tree[C]// 2023 China Automation Congress (CAC). Piscataway: IEEE, 2024: 8648-8652.

[13]

Bhattacharyya R, Wulfe B, Phillips D J, et al. Modeling human driving behavior through generative adversarial imitation learning[J]. IEEE Transactions on Intelligent Transportation Systems, 2023, 24(3): 2874-2887.

[14]

Li Z R, Xiong L, Leng B, et al. Safe reinforcement learning of lane change decision making with risk-fused constraint[C]// 2023 IEEE 26th International Conference on Intelligent Transportation Systems (ITSC). Piscataway: IEEE, 2024: 1313-1319.

[15]

周恒恒, 高松, 王鹏伟, . 基于深度强化学习的智能车辆行为决策研究[J]. 科学技术与工程, 2024, 24(12): 5194-5203.

[16]

Zhou Hengheng, Gao Song, Wang Pengwei, et al. Intelligent vehicles behavior decision-making based on deep reinforcement learning[J]. Science Technology and Engineering, 2024, 24(12): 5194-5203.

[17]

Peng J K, Zhang S Y, Zhou Y, et al. An integrated model for autonomous speed and lane change decision-making based on deep reinforcement learning[J]. IEEE Transactions on Intelligent Transportation Systems, 2022, 23(11): 21848-21860.

[18]

Arjovsky M, Chintala S, Bottou L. Wassertein Generative adversarial networks[C]// Proceedings of the 34th International Conference on Machine Learning. Sydney: PMLR, 2017: 214-223.

[19]

Hidas P. Modelling lane changing and merging in microscopic traffic simulation[J]. Transportation Research Part C: Emerging Technologies, 2002, 10(5/6): 351-371.

[20]

Treiber M, Hennecke A, Helbing D. Congested traffic states in empirical observations and microscopic simulations[J]. Physical Review E, 2000, 62(2): 1805-182

[21]

洪博文, 杜胜品. 智能网联环境下改进的自主换道决策模型[J]. 运筹与模糊学, 2024, 14(6): 607-616.

[22]

Hong Bowen, Du Shengpin. Improved autonomous lane-change decision model in intelligent connected environments[J]. Operations Research and Fuzziology, 2024, 14(6): 607-616.

基金资助

湖北省自然科学基金项目(2023AFB481)

湖北汽车工业学院博士科研启动基金(BK202307)

2025年度湖北省区域科技创新计划国际科技合作项目(2025EHA024)

AI Summary AI Mindmap
PDF (1137KB)

0

访问

0

被引

详细

导航
相关文章

AI思维导图

/