Purposes In recent years, multi-agent systems (MAS) have achieved extensive applications across various domains, particularly in industrial control and automation. However, in partial observability environments, MAS face non-stationarity challenges caused by information incompleteness and dynamic policy interactions, leading to inefficient cooperation and learning convergence difficulties. In this paper, a bidirectional communication-based sequential decision-making model for multi-agent systems is proposed to address these issues through serialized decision processes and bidirectional information exchange. Methods First, traditional parallel decision-making was transformed into a sequential process where agents make decisions successively to reduce policy conflicts. Second, a bidirectional communication module was designed, which integrates forward action intention propagation with reverse attention-driven observation extraction to enhance global perception capabilities. Additionally, a decision scheduling module was introduced to dynamically evaluate agent state values for optimizing decision sequences. Experiments were conducted in the StarCraft Multi-Agent Challenge and Google Research Football environments, covering heterogeneous, homogeneous, and complex adversarial scenarios. Results Results demonstrate that the proposed method significantly outperforms baseline algorithms in win rate and convergence speed, and ablation studies validate the effectiveness of both bidirectional communication and decision scheduling modules.
目前有研究将多智能体并发决策问题转化为序列预测问题,这类方法可以对全局轨迹建模,有效缓解了由于智能体并发决策以及智能体信息缺失引发的非平稳性问题。CHEN et al[11]提出了Decision Transformer方法,该方法将强化学习问题重新定义为序列建模任务,通过序列生成直接建模全局轨迹,避免因智能体策略相互影响导致的动态偏差。MENG et al[12]提出的Offline Pre-trained Multi-Agent Decision Transformer(MADT),利用Transformer架构对智能体的观测序列进行建模,实现了在离线数据上预训练通用策略,通过预训练学习多智能体共享表征,提升对其他智能体策略分布的鲁棒性,从而提升在非平稳环境下的鲁棒性。CAI et al[13]提出的基于Transformer的多智能体强化学习方法COMAT,通过结合图神经网络和自注意力机制,有效捕捉异构多机器人系统中复杂的交互动态,实现了策略在不同团队规模和机器人能力组合上的良好适应性,降低非平稳性干扰。LI et al[14]提出了双向动作依赖协同多智能体Q学习(ACE)方法,该方法通过建模智能体间的双向动作依赖关系,利用协同Q值分解机制来优化联合策略,从而提升多智能体协作效率。其优点是能够有效捕捉智能体间的复杂交互,并在非全知环境下实现更稳定的策略学习。
此外本文在GRF(Google Research Football)环境进行相关实验。与SMAC相比,GRF 2020提供了一个更具挑战性的环境,动作空间更大且奖励稀疏。在GRF中,智能体需要协调时间和位置来组织进攻,只有得分才能获得奖励。在本文实验中,智能体控制左队除守门员之外的球员,而游戏内置引擎控制右队球员。本文在两个具有挑战性的场景中进行实验:academy_3_vs_1_with_keeper, academy_counterattack_hard[16]。
FANC, ZENGL, SUNY,et al.Finding key players in complex networks through deep reinforcement learning[J].Nature Machine Intelligence,2020,2(6):317-324.
[2]
DEGRAVEJ, FELICIF, BUCHLIJ,et al.Magnetic control of tokamak plasmas through deep reinforcement learning[J].Nature,2022,602(7897):414-419.
[3]
LOWER, WUY, TAMARA,et al.Multi-agent actor-critic for mixed cooperative-competitive environments[C]//Advances in Neural Information Processing Systems (NeurIPS).Long Beach,USA:Carran Associates Inc,2017:6379-6390.
[4]
HERNÁNDEZ-LEALP, KARTALB, TAYLORM E.A survey and critique of multi-agent deep reinforcement learning[J].Autonomous Agents and Multi-Agent Systems,2019,33(6):750-797.
[5]
PAPOUDAKISG, CHRISTIANOSF, RAHMANA,et al.Dealing with non-stationarity in multi-agent deep reinforcement learning[EB/OL].arXiv preprint arXiv:
[6]
FOERSTERJ, NARDELLIN, FARQUHARG,et al.Stabilising experience replay for deep multi-agent reinforcement learning[C]//Proceedings of the 34th International Conference on Machine Learning,2017:1146-1155.
[7]
YUNW J,LIM B, JUNGS,et al.Attention-based reinforcement learning for rcaltime UAV semantic communication[C]//2021 17th International Symposium on wireless Commuincation System(ISWCS).Berlin,Germany:IEEE,2021:1-6.
[8]
CHARAPKOA, AILIJIANGA, DEMIRBASM.Pigpaxos: Devouring the communication bottlenecks in distributed consensus[C]//Proceedings of the 2021 International Conference on Management of Data.Visual Event:ACM,2021: 235-247.
[9]
SUNC, ZANGZ, LIJ,et al.T2MAC: Targeted and trusted multi-agent communication through selective engagement and evidence-driven integration[C]//Proceedings of the AAAI Conference on Artificial Intelligence.Vancouver,Canada:AAAI Press,2024,38(13),15154-15163.
[10]
YUC, VELUA, VINITSKYE,et al.The surprising effectiveness of PPO in cooperative multi-agent games[J].Advances in Neural Information Processing Systems,2022,35:24611-24624.
[11]
CHENL, LUK, RAJESWARANA,et al.Decision transformer:Reinforcement learning via sequence modeling[J].Advances in Neural Information Processing Systems,2021,34:15084-15097.
CAIY, HEX, GUOH,et al.Transformer-based multi-agent reinforcement learning for generalization of heterogeneous multi-robot cooperation[C]//Proceeding of the IEEE/RSJ International Conference on Intellgent Robots and systems.Abu Dhabi,United Arab Fmirates:IEEE,2014:13695-13702.
[14]
LIC, LIUJ, ZHANGY,et al.ACE: Cooperative multi-agent Q-learning with bidirectional action-dependency[C]//Proceedings of the AAAI Conference on Artificial Intelligence.Washinton,DC:AAAI Press,2023,37(7),8536-8544.
[15]
OLIEHOEKF A, AMATOC.A concise introduction to decentralized POMDPs.Springer International Publishing[M].Cham:Springer Interctional Publishing,2016.
[16]
KURACHK, RAICHUKA, STAŃCZYKP,et al.Google research football:A novel reinforcement learning environment[C]//Proceedings of the AAAI Conference on Artificial Intelligence.New York:NY:AAAI Press,2020,34(4):4501-4510.
[17]
CONTRIBUTORSOPENDILAB.DI-engine: OpenDILab decision intelligence engine[C]//Proceedings of the AAAI Conference on Artificial Intelligence,Virtual Event:AAAI Press,2021,35(13),11956-11964.
[18]
RASHIDT, SAMVELYANM, DE WITTC S,et al.QMIX: Monotonic value function factorisation for deep multi-agent reinforcement learning[J].Journal of Machine learning Research,2020,21(178):1-51.