1.MIIT Key Laboratory of Pattern Analysis and Machine Intelligence, College of Computer Science and Technology, Nanjing University of Aeronautics and Astronautics,Nanjing,211106,China
2.School of Software,Nanjing University,Nanjing,210093,China
Efficiently leveraging imperfect expert demonstration data is one of the key challenges in the field of imitation learning. Imperfect demonstrations,which entail the loss of fine⁃grained local details,can lead to error accumulation. To address scenarios with incomplete demonstrations,we propose a robust imitation learning method based on generative trajectory modeling. Our approach utilizes a Decision Transformer to generate continuous state⁃action sequences conditioned on incomplete demonstrations,thereby completing the trajectories. The quality of these generated trajectories is controlled by setting a target return. The completed trajectories are then fed into a diffusion model for further trajectory generation. Additionally,we introduce a novel scoring mechanism to evaluate whether the trajectories generated by the diffusion model conform to the environmental dynamics constraints present in the expert demonstrations. This model can generate high⁃reward trajectories that adhere to the dynamics constraints while effectively completing imperfect trajectories,thereby significantly reducing error accumulation and improving planning accuracy. Experiments show that our method outperforms existing offline reinforcement learning approaches. It successfully addresses the planning⁃to⁃reality mismatch problem in high⁃dimensional continuous spaces and demonstrates superior robustness and generalization capabilities.
Wang et al[8]的扩散净化模仿学习(Diffusion Purified Imitation Learning,DP⁃IL),使用扩散模型对次优演示进行“前向加噪⁃反向去噪”的两阶段净化,使数据分布回到接近最优的情况,有效缓解演示噪声问题,在多个控制任务中取得显著性能提升,但需依赖少量最优示范进行建模且计算开销较高.Wu et al[9]的2IWIL (Two⁃step Importance Weighting Imitation Learning)与IC⁃GAIL (Generative Adversarial Imitation Learning with Imperfect Demonstration and Confidence)两种方法,通过部分置信度标注进行模仿学习,即分别使用重要性加权和对抗学习机制从不完美数据中恢复最优策略,理论上提供了泛化误差界,但需要额外的置信度标注,限制了其实用性.Xu et al[10]的判别器加权行为克隆(Discriminator⁃Weighted Behavioral Cloning,DWBC),引入判别器来区分专家与非专家数据,并以最坏情况误差最小化为目标来提升策略鲁棒性,避免了复杂的对抗训练,具有较高效率,但仍依赖数据分布覆盖,难以应对严重分布偏移.
ChenL L, LuK, RajeswaranA,et al. Decision transformer:Reinforcement learning via sequence modeling∥Proceedings of the 35th International Conference on Neural Information Processing Systems. Red Hook,NY,USA:Curran Associates Inc,2021:15084-15097.
[3]
HoJ, JainA, AbbeelP. Denoising diffusion probabilistic models. Advances in Neural Information Processing Systems,2020,33:6840-6851.
[4]
CroitoruF A, HondruV, IonescuR T,et al. Diffusion models in vision: A survey. IEEE Transactions on Pattern Analysis and Machine Intelligence,2023,45(9):10850-10869.
[5]
JannerM, DuY L, TenenbaumJ B,et al. Planning with diffusion for flexible behavior synthesis∥Proceedings of the 39th International Conference on Machine Learning. New York,NY,USA:PMLR,2022:9902-9915.
[6]
HuangK, ScaliseR, WinstonC,et al. Using non⁃expert data to robustify imitation learning via offline reinforcement learning. https://arxiv.org/abs/2510.19495,2025-10-25.
[7]
YuX R, HanB, TsangI W. USN:A robust imitation learning method against diverse action noise. Journal of Artificial Intelligence Research,2024,79:1237-1280.
XuH R, ZhanX Y, YinH L,et al. Discriminator⁃weighted offline imitation learning from suboptimal demonstrations. https://arxiv.org/abs/2207.10050,2022-07-20.
[11]
JiangK, YaoJ Y, TanX Y,et al. Recovering from out⁃of⁃sample states via inverse dynamics in offline reinforcement learning. Advances in Neural Information Processing Systems,2023,36:38815-38826.
[12]
LevineS, KumarA, TuckerG,et al. Offline reinforcement learning:Tutorial,review,and perspectives on open problems. https://arxiv.org/abs/2005.01643,2020-11-01.
VaswaniA, ShazeerN, ParmarN,et al. Attention is all you need∥Proceedings of the 31st International Conference on Neural Information Processing Systems.Red Hook,NY,USA:Curran Associates Inc.,2017:6000-6010.
[15]
Sohl⁃DicksteinJ, WeissE A, MaheswaranathanN,et al. Deep unsupervised learning using nonequilibrium thermodynamics. https://arxiv.org/abs/1503.03585,2015-11-18.
[16]
LevineS. Reinforcement learning and control as probabilistic inference:Tutorial and review. https://arxiv.org/abs/1805.00909,2018-05-20.
[17]
FuJ, KumarA, NachumO,et al. D4RL:Datasets for deep data⁃driven reinforcement learning. https://arxiv.org/abs/2004.07219,2021-02-06.