The evolution of end-to-end autonomous driving is systematically reviewed, tracing the progression from modular pipelines to fully end-to-end architectures. Representative models and core technologies at each developmental stage are analyzed, and the causal relationships underlying paradigm shifts are uncovered. In particular, the integration of large-scale models with autonomous driving systems is highlighted. In addition, a comprehensive overview of industrial applications and future prospects of autonomous driving is provided, along with a detailed classification and evaluation of relevant datasets and an exploration of future challenges in end-to-end autonomous driving. A holistic perspective on the current state of end-to-end autonomous driving and directions for future research is thereby offered to the research community.
目前已有多篇英文综述系统总结了端到端自动驾驶的发展趋势、挑战与方法,例如,Chen等[1]从模仿学习和强化学习的视角系统梳理了端到端模型的发展脉络;Singh[2]研究了完全可微的端到端强化学习技术。Yang等[3]研究了大语言模型(Large language model, LLM)在AD中的使用情况,深入分析了各类别代表性方法及其相关数据集。Liu等[4]从多个角度对265个自动驾驶数据集进行综合评估。然而,当前综述往往偏向特定算法或局部技术视角,且中文领域的端到端自动驾驶综述相对匮乏,且现有研究多集中于单一模块讨论(如感知[5]或决策[6]),未能从整体架构演进的角度系统揭示端到端方法的技术驱动力与发展逻辑[7]。褚端峰等[8]对生成式人工智能在端到端方法中的应用进行了初步探讨,但工作往往缺乏对大型多模态模型在自动驾驶中的作用、不同范式之间的互补关系以及新型生成式模型的系统梳理。针对以上不足,本文从整体技术演化的视角出发,全面回顾端到端自动驾驶技术的发展历程,深入分析其各阶段的核心技术驱动力,并且在此基础上提出补充性视角,并以中文形式组织,亦有利于加快中文研究群体的信息获取。本文重点探讨了大模型在端到端方法中的应用潜力与发展方向,并系统总结其在感知、决策等关键任务中面临的技术挑战。本文贡献如下:
这一阶段从基于仅处理语言信息的LLM,逐步发展到基于可处理视觉语言等多模态的大模型(Multimodal large language model, MLLM),来生成粗指令辅助自动驾驶决策。随着多模态数据与动作编码的引入,可直接生成自动驾驶动作序列的视觉语言动作模型(Vision language action, VLA)成为新的范式,实现从感知到控制的真正端到端模型,如图2所示,输入内容从局部端到端的图片视频扩展到包含文本信息的大模型完全端到端,而基于VLA则在此基础上加入历史动作序列、实时动作反馈、物理交互数据等,构建更完整的物理世界表征。输出也从使用基于规则的控制器到预测离散轨迹序列的回归模型再到使用从纯噪声中逐步恢复轨迹的条件扩散模型。
ChenL, WuP, ChittaK, et al. End‑to‑end autonomous driving: Challenges and frontiers[J]. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2024, 46(12): 10164-10183.
[2]
SinghA. End-to-end autonomous driving using deep learning: a systematic review[J/OL]. [2023-11-30].
[3]
YangZ, JiaX, LiH, et al. LLM4Drive: a survey of large language models for autonomous driving[J/OL]. [2025-09-15].
[4]
LiuM, YurtseverE, FossaertJ, et al. A survey on autonomous driving datasets: statistics, annotation quality, and a future outlook[J]. IEEE Transactions on Intelligent Vehicles, 2024, 9(3): 1024-1045.
HuangDe-qi, HuangHai-feng, HuangDe-yi, et al. Survey on the application of BEV perception learning in autonomous driving[J]. Computer Engineering and Applications, 2025, 61(2): 1-23.
LiuYi-fei, HuXue-min, ChenGuo-wen, et al. A survey on end-to-end autonomous driving motion planning based on visual perception[J]. Journal of Image and Graphics, 2021, 26(1): 49-66.
ChenYan-yan, TianDa-xin, LinChun-mian, et al. Survey of end-to-end autonomous driving systems[J]. Journal of Image and Graphics, 2024, 29(11): 3216-3237.
ChuDuan-feng, WangRu-kang, WangJing-yi, et al. Research progress and challenges in end-to-end autonomous driving[J]. China Journal of Highway and Transport, 2024, 37(10): 209-232.
[13]
GaoH, LiY, LongK, et al. A survey for foundation models in autonomous driving[J/OL]. [2024-02-02].
[14]
TampuuA, MatiisenT, SemikinM, et al. A survey of end-to-end driving: architectures and training methods[J]. IEEE Transactions on Neural Networks and Learning Systems, 2020, 33(7): 1364-1384.
[15]
TengS, HuX, DengP, et al. Motion planning for autonomous driving: the state of the art and future perspectives[J]. IEEE Transactions on Intelligent Vehicles, 2023, 8(6): 3692-3711.
[16]
FouratiS, JaafarW, BaccarN, et al. XLM for autonomous driving systems: a comprehensive review[J/OL]. [2024-09-16].
[17]
ShiJ, ChenJ, WangY, et al. Motion forecasting for autonomous vehicles: a survey[C]∥Proceedings of the 2025 IEEE International Conference on Robotics and Automation (ICRA), Seattle, USA, 2025: 1-10.
[18]
FengT, WangW, YangY. A survey of world models for autonomous driving[J/OL]. [2025-06-19].
[19]
DalalN, TriggsB. Histograms of oriented gradients for human detection[C]∥Proceedings of the 2005 IEEE Computer Society Conference on Computer Vision and Pattern Recognition(CVPR), San Diego, USA, 2005: 886-893.
[20]
BansalM, KrizhevskyA, OgaleA S. ChauffeurNet: Learning to drive by imitating the best and synthesizing the worst[J/OL]. [2018-12-07].
[21]
VaswaniA, ShazeerN, ParmarN, et al. Attention is all you need[C]∥Proceedings of the 31st Conference on Neural Information Processing Systems(NeurIPS 2017), Long Beach, USA: NeurIPS Foundation, 2017: 5998-6008.
[22]
ZhouB, KrahenbuhlP. Cross-view transformers for real-time map-view semantic segmentation[C]∥Proceedings of the 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition(CVPR), New Orleans, USA: IEEE, 2022: 13750-13759.
[23]
XuR, TuZ, XiangH, et al. CoBEVT: Cooperative bird's eye view semantic segmentation with sparse transformers[J/OL]. [2022-07-05].
[24]
ZhangY, ZhuZ, DuD. OccFormer: Dual-path transformer for vision-based 3D semantic occupancy prediction[C]∥Proceedings of the 2023 IEEE/CVF International Conference on Computer Vision (ICCV), Paris, France, 2023: 9399-9409.
[25]
WangX, ZhuZ, XuW, et al. OpenOccupancy: A large scale benchmark for surrounding semantic occupancy perception[C]∥Proceedings of the 2023 IEEE/CVF International Conference on Computer Vision (ICCV), Paris, France, 2023: 17804-17813.
[26]
FanH, ZhuF, LiuC, et al. Baidu Apollo EM motion planner[J/OL]. [2018-07-20].
WangXiao, ZhangXiang-yu, ZhouRui, et al. Research on cognitive autonomous driving intelligent architecture based on parallel testing[J]. Acta Automatica Sinica, 2024, 50(2): 356-371.
[29]
MnihV, KavukcuogluK, SilverD, et al. Human-level control through deep reinforcement learning[J]. Nature, 2015, 518(7540): 529-533.
[30]
RezaeeK, YadmellatP, NosratiM S, et al. Multi-lane cruising using hierarchical planning and reinforcement learning[C]∥Proceedings of the IEEE Intelligent Transportation Systems Conference(ITSC), Auckland, New Zealand, 2019: 1800-1806.
[31]
WolfP, KurzerK, WingertT, et al. Adaptive behavior generation for autonomous driving using deep reinforcement learning with compact semantic states[C]∥Proceedings of the 2018 IEEE Intelligent Vehicles Symposium(IV), Changshu, China, 2018: 993-1000.
[32]
WangP, LiuD, ChenJ, et al. Human-like decision making for autonomous driving via adversarial inverse reinforcement learning[J/OL]. [2019-11-19].
[33]
SunL, PengC-T J, ZhanW, et al. A fast integrated planning and control framework for autonomous driving via imitation learning[J/OL]. [2017-07-09].
[34]
WulfmeierM, RaoD, WangD Z, et al. Large-scale cost function learning for path planning using deep inverse reinforcement learning[J]. The International Journal of Robotics Research, 2017, 36(10): 1073-1087.
[35]
HausknechtM J, StoneP. Deep reinforcement learning in parameterized action space[J/OL]. [2015-11-13].
[36]
ChenJ, WangZ, TomizukaM. Deep hierarchical reinforcement learning for autonomous driving with distinct behaviors[C]∥Proceedings of the 2018 IEEE Intelligent Vehicles Symposium (IV), Changshu, China: IEEE, 2018: 1239-1244.
[37]
ChenD, ZhouB, KoltunV, et al. Learning by cheating[J/OL]. [2019-12-27].
[38]
ChenD, KrähenbühlP. Learning from all vehicles[C]∥Proceedings of the 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), New Orleans, USA: IEEE, 2022: 17201-17210.
[39]
RusuA A, ColmenarejoS G, GülçehreÇ, et al. Policy distillation[J/OL]. [2015-11-19].
[40]
HouY, MaZ, LiuC, et al. Learning to steer by mimicking features from heterogeneous auxiliary networks[C]∥Proceedings of the 32nd AAAI Conference on Artificial Intelligence, New Orleans, USA, 2018: 1-9.
[41]
ZhaoA, HeT, LiangY, et al. SAM: Squeeze-and-mimic networks for conditional visual driving policy learning[C]∥Proceedings of the 2019 Conference on Robot Learning, Osaka, Japan: PMLR, 2019: 1-10.
[42]
ZhangZ, LinigerA, DaiD, et al. End-to-end urban driving by imitating a reinforcement learning coach[C]∥Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), Montreal, Canada: IEEE, 2021: 15202-15212.
[43]
JiaX, GaoY, ChenL, et al. DriveAdapter: Breaking the coupling barrier of perception and planning in end-to-end autonomous driving[C]∥Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), Paris, France, 2023: 7919-7929.
[44]
SharifzadehS, ChiotellisI, TriebelR, et al. Learning to drive using inverse reinforcement learning and deep Q-networks[J/OL].
[45]
AlizadehA, MoghadamM, BicerY, et al. Automated lane change decision making using deep reinforcement learning in dynamic and uncertain highway environment[C]∥Proceedings of the IEEE Intelligent Transportation Systems Conference(ITSC), Auckland, New Zealand: IEEE, 2019: 1399-1404.
[46]
ZhangJ, HuangZ, Ohn-BarE. Coaching a teachable student[C]∥Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Vancouver, Canada, 2023: 7805-7815.
[47]
ChenD, KoltunV, KrähenbühlP. Learning to drive from a world on rails[C]∥Proceedings of the 2021 IEEE/CVF International Conference on Computer Vision(ICCV), Montreal, Canada, 2021: 15570-15579.
[48]
ChenJ, YuanB, TomizukaM. Deep imitation learning for autonomous driving in generic urban scenarios with enhanced safety[C]∥Proceedings of the 2019 IEEE/RSJ International Conference on Intelligent Robots and Systems(IROS), Macau, China, 2019: 2884-2890.
[49]
RossS, GordonG J, BagnellJ A. A reduction of imitation learning and structured prediction to no-regret online learning[C]∥Proceedings of the 13th International Conference on Artificial Intelligence and Statistics, Sardinia, Italy, 2010: 627-635.
[50]
ChittaK, PrakashA, JaegerB, et al. TransFuser: Imitation with transformer-based sensor fusion for autonomous driving[J]. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2022, 45(11): 12878-12895.
[51]
SadatA, CasasS, RenM, et al. Perceive, predict, and plan: Safe motion planning through interpretable semantic representations[J/OL].
[52]
CasasS, SadatA, UrtasunR. MP3: A unified model to map, perceive, predict and plan[C]∥Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition(CVPR), Nashville, USA: IEEE, 2021: 14398-14407.
[53]
HuS, ChenL, WuP, et al. ST-P3: End-to-end vision-based autonomous driving via spatial-temporal feature learning[C]∥Proceedings of the 2022 European Conference on Computer Vision(ECCV), Tel Aviv, Israel, 2022: 1-17.
[54]
HuY, YangJ, ChenL, et al. Planning-oriented autonomous driving[C]∥Proceedings of the 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition(CVPR), Vancouver, Canada, 2023: 17853-17862.
[55]
JiangB, ChenS, XuQ, et al. VAD: Vectorized scene representation for efficient autonomous driving[C]∥Proceedings of the 2023 IEEE/CVF International Conference on Computer Vision(ICCV), Paris, France, 2023: 8306-8316.
[56]
CHENS Y, JIANGB, GAOH, et al. VADv2: End-to-end vectorized autonomous driving via probabilistic planning[J/OL].
[57]
SUNW C, LINX W, SHIY N, et al. SparseDrive: End-to-end autonomous driving via sparse scene representation[J/OL]. [2024-05-30].
[58]
Johnson-RobersonM, BartoC, MehtaR, et al. Driving in the matrix: Can virtual worlds replace human-generated annotations for real world tasks?[C]∥Proceedings of the IEEE International Conference on Robotics and Automation (ICRA), Singapore, 2017: 746-753.
[59]
XuZ, ZhangY, XieE, et al. DriveGPT4: Interpretable end-to-end autonomous driving via large language model[J]. IEEE Robotics and Automation Letters, 2023, 9(9): 8186-8193.
[60]
TouvronH, LavrilT, IzacardG, et al. LLaMA: Open and efficient foundation language models[J/OL]. [2023-02-27].
[61]
GunasekarS, ZhangY, AnejaJ, et al. Textbooks are all you need[J/OL]. [2023-06-20].
[62]
RadfordA, KimJ W, HallacyC, et al. Learning transferable visual models from natural language supervision[C]∥Proceedings of the 38th International Conference on Machine Learning(ICML 2021), Virtual Event: PMLR, 2021: 8748-8763.
[63]
JiaC, YangY, XiaY, et al. Scaling up visual and vision-language representation learning with noisy text supervision[C]∥Proceedings of the 38th International Conference on Machine Learning(ICML 2021), Online, 2021: 4904-4916.
[64]
BaiJ, BaiS, ChuY, et al. Qwen Technical Report[J/OL]. [2023-09-28].
[65]
Deepseek-AI. DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning[R]. Beijing: DeepSeek Inc., 2025.
[66]
AlayracJ-B, DonahueJ, LucP, et al. Flamingo: A visual language model for few-shot learning[J/OL]. [2022-04-29].
[67]
ScaoT L, FanA, AkikiC, et al. BLOOM: A 176B-parameter open-access multilingual language model[J/OL]. [2022-11-09].
[68]
DevlinJ, ChangM-W, LeeK, et al. BERT: Pre-training of deep bidirectional transformers for language understanding[C]∥Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies(NAACL-HLT), Minneapolis, USA, 2019: 4171-4186.
[69]
ShaH, MuY, JiangY, et al. LanguageMPC: Large language models as decision makers for autonomous driving[J/OL]. [2023-10-04].
[70]
PengM, GuoX, ChenX, et al. LC-LLM: Explainable lane-change intention and trajectory predictions with large language models[J/OL]. [2024-03-27].
[71]
WangT, XieE, ChuR, et al. DriveCoT: Integrating chain-of-thought reasoning with end-to-end driving[J/OL].
[72]
CaesarH, KabzanJ, TanK S, et al. nuPlan: A closed-loop ML-based planning benchmark for autonomous vehicles[J/OL]. [2021-06-22].
[73]
LiJ, LiD, XiongC, et al. BLIP: bootstrapping language-image pre-training for unified vision-language understanding and generation[C]∥Proceedings of the 39th International Conference on Machine Learning (ICML 2022),Baltimore, USA, 2022: 12888-12900.
[74]
WuD, HanW, WangT, et al. Language prompt for autonomous driving[J/OL]. [2023-09-08].
[75]
DingX, HanJ, XuH, et al. HiLM-D: Towards high-resolution understanding in multimodal large language models for autonomous driving[J/OL]. [2023-09-11].
[76]
LiangM, SuJ-C, SchulterS, et al. AIDE: An automatic data engine for object detection in autonomous driving[C]∥Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Seattle, USA, 2024: 14695-14706.
[77]
WangW, XieJ, HuC, et al. DriveMLM: Aligning multi-modal large language models with behavioral planning states for autonomous driving[J/OL]. [2023-12-14].
[78]
KimM J, PertschK, KaramchetiS, et al. OpenVLA: an open-source vision-language-action model[R]. San Francisco: Stanford University, 2024.
[79]
XingS, QianC Y, WangY P, et al. OpenEMMA: open-source multimodal model for end-to-end autonomous driving[J/OL]. [2024-12-19].
[80]
AraiH, MiwaK, SasakiK, et al. CoVLA: Comprehensive vision-language-action dataset for autonomous driving[J/OL]. [2024-08-19].
[81]
Sohl-DicksteinJ N, WeissE A, MaheswaranathanN, et al. Deep unsupervised learning using nonequilibrium thermodynamics[J/OL].
[82]
WenJ J, ZhuM J, ZhuY C, et al. Diffusion-VLA: Generalizable and interpretable robot foundation model via self-generated reasoning[J/OL]. [2024-12-04].
[83]
ZhengY, LiangR, ZhengK, et al. Diffusion-based planning for autonomous driving with flexible guidance[J]. [2025-01-26].
[84]
LiaoB, ChenS, YinH, et al. Diffusiondrive: truncated diffusion model for end-to-end autonomous driving[J]. [2024-11-22].
[85]
DawidA, LeCunY. Introduction to latent variable energy-based models: a path towards autonomous machine intelligence[J/OL]. [2023-06-05].
[86]
KimS W, PhilionJ, TorralbaA, et al. DriveGAN: towards a controllable high-quality neural simulation[C]∥Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition(CVPR), Nashville, USA, 2021: 5816-5825.
[87]
HuA, RussellL, YeoH, et al. Gaia-1: a generative world model for autonomous driving[J/OL]. [2023-09-29].
[88]
HaD, SchmidhuberJ. Recurrent world models facilitate policy evolution[C]∥Proceedings of the 32nd Conference on Neural Information Processing Systems(NeurIPS 2018), Montreal, Canada, 2018: 1-10.
[89]
HafnerD, LillicrapT, BaJ, et al. Dream to control: Learning behaviors by latent imagination[J/OL]. [2019-12-04].
[90]
PanM, ZhuX, WangY, et al. Iso-Dream: Isolating and leveraging noncontrollable visual dynamics in world models[C]∥Proceedings of the 36th Conference on Neural Information Processing Systems (NeurIPS 2022), OrleansNew, USA, 2022: 1-12.
[91]
JiaF, MaoW, LiuY, et al. Adriver-i: a general world model for autonomous driving[J/OL]. [2023-11-22].
[92]
CenJ, YuC H, YuanH J, et al. WorldVLA: Towards autoregressive action world model[J/OL]. [2023-11-22].
[93]
KimJ, RohrbachA, DarrellT, et al. Textual explanations for self-driving vehicles[C]∥Proceedings of the 15th European Conference on Computer Vision (ECCV 2018), Munich, Germany, 2018: 577-593.
[94]
XuY, YangX, GongL, et al. Explainable object-induced action decision for autonomous vehicles[C]∥Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition(CVPR), Seattle, USA, 2020: 9520-9529.
[95]
Ben-YounesH, ZablockiÉ, PérezP, et al. Driving behavior explanation with multi-level fusion[J/OL]. [2020-12-09].
[96]
BrooksT, PeeblesB, HolmesC, et al. Video generation models as world simulators[J/OL]. [2024-02-22].
[97]
BinasJ, NeilD, LiuS-C, et al. DDD17: end-to-end DAVIS driving dataset[J/OL]. [2017-11-03].
[98]
MaddernW P, PascoeG, LinegarC, et al. 1 year, 1000 km: the Oxford RobotCar dataset[J]. The International Journal of Robotics Research, 2017, 36(1): 15-31.
[99]
DosovitskiyA, RosG, CodevillaF, et al. CARLA: an open urban driving simulator[C]∥Proceedings of the 1st Conference on Robot Learning(CoRL 2017. ViewMountain, USA, 2017: 1-16.
[100]
HeckerS, DaiD, GoolL V. End-to-end learning of driving models with surround-view cameras and route planners[C]∥Proceedings of the 15th European Conference on Computer Vision(ECCV 2018), Munich, Germany, 2018: 435-453.
[101]
RamanishkaV, ChenY-T, MisuT, et al. Toward driving scene understanding: a dataset for learning driver behavior and causal reasoning[C]∥Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition(CVPR),Salt Lake City, USA, 2018: 7699-7707.
[102]
SunP, KretzschmarH, DotiwallaX, et al. Scalability in perception for autonomous driving: Waymo Open dataset[C]∥Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition(CVPR), Seattle, USA, 2020: 2443-2451.
[103]
CaesarH, BankitiV, LangA H, et al. nuScenes: a multimodal dataset for autonomous driving[C]∥Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition(CVPR), Seattle, USA, 2020: 11618-11628.
[104]
ChenY, WangJ, LiJ, et al. LiDAR-Video driving dataset: Learning driving policies effectively[C]∥Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition(CVPR), Salt Lake City, USA, 2018: 5870-5878.
[105]
GeyerJ, KassahunY, MahmudiM, et al. A2D2: Audi autonomous driving dataset[J/OL]. [2020-04-14].
[106]
DeruyttereT, VandenhendeS, GrujicicD, et al. Talk2Car: Taking control of your self-driving car[J/OL]. [2019-09-24].
[107]
MallaS, ChoiC, DwivediI, et al. DRAMA: joint risk localization and captioning in driving[C]∥Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision(WACV), Waikoloa, USA, 2023: 1043-1052.
[108]
SimaC, RenzK, ChittaK, et al. DriveLM: Driving with graph visual question answering[J/OL]. [2023-12-21].
[109]
QianT, ChenJ, ZhuoL, et al. nuScenes-QA: a multi-modal visual question answering benchmark for autonomous driving scenario[J/OL]. [2023-05-24].
[110]
FangJ, LiL, YangK, et al. Cognitive accident prediction in driving scenes: a multimodality benchmark[J/OL]. [2022-12-19].
[111]
ParkS, LeeM, KangJ, et al. VLAAD: Vision and language assistant for autonomous driving[C]∥Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision Workshops (WACVW), Seattle, USA, 2024: 980-987.
[112]
NiZ, DuS, HouZ, et al. Para-Lane: Multi-lane dataset registering parallel scans for benchmarking novel view synthesis[R]. Beijing: Tsinghua University, 2025.
[113]
Momenta. Momenta announces new autonomous driving solution powered by NVIDIA DRIVE orin, commercializing urban NOA on scale[EB/OL]. [2024-04-26].
[114]
Huawei Technologies Co., Ltd. Striding towards the intelligent world: the journey to all intelligence 2024[EB/OL]. [2025-10-01].
[115]
Amara. Top stories of Li Auto in 2024[EB/OL]. [2025-02-03].
[116]
Gabriella. Baidu Apollo launches Apollo ADFM large model, sixth-gen unmanned Robotaxi at Apollo Day 2024[EB/OL]. [2024-05-15].
[117]
ZhaoG, NiC, WangX, et al. DriveDreamer 4D: world models are effective data machines for 4D driving scene representation[J/OL]. [2024-03-04].
[118]
XuT, ChenZ, WuL, et al. Motion dreamer: boundary conditional motion reasoning for physically coherent video generation[J/OL]. [2023-11-25].
[119]
KimS W, PhilionJ, TorralbaA, et al. DriveGAN: towards a controllable high-quality neural simulation[C]∥Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Online,2021: 1231-1240.
[120]
ZhouX, LinZ, ShanX, et al. Driving Gaussian: composite gaussian splatting for surrounding dynamic autonomous driving scenes[J/OL]. [2024-03-15].
[121]
ZablockiÉ, Ben-YounesH, PérezP, et al. Explainability of deep vision-based autonomous driving systems: Review and challenges[J]. International Journal of Computer Vision, 2021, 130(10): 2425-2452.
[122]
MahawattaM A, Cabrero-DanielB, YuY, et al. LLMs can check their own results to mitigate hallucinations in traffic understanding tasks[J/OL]. [2024-09-19].
[123]
FanJ, WuJ, ChuH,et al. Hallucination elimination and semantic enhancement framework for vision-language models in traffic scenarios[J/OL]. [2024-12-11].
[124]
WangJ, WuZ, DongQ, et al. Hybrid-Driving: an autonomous driving decision framework integrating large language models, knowledge graphs and driving rules[J]. Proceedings of the AAAI Conference on Artificial Intelligence, 2025, 39(1): 826-833.
[125]
ShaoH, WangL, ChenR, et al. Safety-enhanced autonomous driving using interpretable sensor fusion transformer[C]∥Proceedings of the Conference on Robot Learning(CoRL), Auckland, New Zealand, 2022.
[126]
ChittaK, PrakashA, GeigerA. NEAT: Neural attention fields for end-to-end autonomous driving[C]∥Proceedings of the IEEE/CVF International Conference on Computer Vision(ICCV), Montreal, Canada, 2021: 15773-15783.
[127]
RuanK, DiX. Learning human driving behaviors with sequential causal imitation learning[J]. Proceedings of the AAAI Conference on Artificial Intelligence, 2022, 36(12): 4583-4590.
[128]
WangB, LiaoH, WangC, et al. Beyond patterns: Harnessing causal logic for autonomous driving trajectory prediction[C]∥Proceedings of the Thirty-Third International Joint Conference on Artificial Intelligence(IJCAI-24), Jeju, South Korea, 2024: 3124–3132.
[129]
JiangW, WangL, ZhangT, et al. RobustE2E: exploring the robustness of end-to-end autonomous driving[J]. Electronics, 2024, 13(16): No.3299.
[130]
SathyamR, LiY. Foundation models for autonomous driving perception: a survey through core capabilities[J]. IEEE Open Journal of Vehicular Technology, 2025,6: 2554-2582.
[131]
GaoZ, MuY, ShenR, et al. Enhance sample efficiency and robustness of end-to-end urban autonomous driving via semantic masked world model[J]. IEEE Transactions on Intelligent Transportation Systems, 2022, 25(12): 13067-13079.
[132]
TonevaM, SordoniA, CombesR T D, et al. An empirical study of example forgetting during deep neural network learning[J/OL]. [2018-12-12].
[133]
KhoslaS, ZhuZ, HeY. Survey on memory-augmented neural networks: cognitive insights to AI applications[J/OL]. [2023-12-11].
[134]
DohareS, Hernandez-GarciaJ F, LanQ, et al. Loss of plasticity in deep continual learning[J]. Nature, 2024, 632(768): 768-774.
[135]
KaiwartyaO, AbdullahA H, CaoY, et al. Internet of vehicles: motivation, layered architecture, network model, challenges, and future aspects[J]. IEEE Access, 2016, 4: 5356-5373.
[136]
FengD, Haase-SchützC, RosenbaumL, et al. Deep multi-modal object detection and semantic segmentation for autonomous driving: Datasets, methods, and challenges[J]. IEEE Transactions on Intelligent Transportation Systems, 2019, 22(3): 1341-1360.
[137]
NguyenA, DoT K L, TranM-N, et al. Deep federated learning for autonomous driving[C]∥Proceedings of the 2022 IEEE Intelligent Vehicles Symposium(IV), Aachen, Germany, 2022: 1824-1830.