arXiv AI By Jinyuan Zu, Xiaowei Lv, Yongcai Wang, Deying Li, Yunjun Han, Wenping Chen, Fengyi Zhang, Naiqi Wu

CCKS: Consensus-based Communication and Knowledge Sharing

Read the original on arXiv AI →

arXiv:2606. 12281v1 Announce Type: cross Abstract: In Decentralized Training and Decentralized Execution (DTDE) for cooperative Multi-Agent Reinforcement Learning (MARL), action-advising-based knowledge sharing promotes interpretable and scalable cooperation among agents.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Machine Learning
Aug 6

Communication-Enhanced Tutoring for Efficient Decentralized Multi-Agent Reinforcement Learning

arXiv:2508. 13661v4 Announce Type: replace Abstract: Centralized Training with Decentralized Execution (CTDE) is the dominant paradigm in multi-agent reinforcement learning (MARL), enabling agents to act independently at test time while leveraging additional information during training.

By Maciej Wojtala, Bogusz Stefa\'nczyk, Dominik Bogucki, {\L}ukasz Lepak, Pawe{\l} Wawrzy\'nski
arXiv AI
6d ago

G2MAF: Test-Time Gradient Guidance for Multi-Agent Flow Policies

G2MAF is a test‑time refinement framework for offline multi‑agent reinforcement learning that applies a single globally normalized, projected critic gradient to adjust all agents’ actions while keeping them close to a frozen policy proposal. The method improves performance on 24 Multi‑Party Environment (MPE) and StarCraft Multi‑Agent Challenge (SMAC) benchmarks, achieving mean relative gains of 9.2% on MPE and 8.9% on SMAC, with only a 6% increase in inference latency.

By Guowei Zou, Haitao Wang, Guoxin Wang, Zhiquan Chen, Beiwen Zhang, Guojie Wang, Hejun Wu
arXiv AI
Aug 28

SIGMA: Structured Noise-Effect-Aware Grouped Multi-Agent Aggregation

SIGMA is a hierarchical framework for cooperative multi‑agent reinforcement learning that addresses structured noise effects—local correlations in noise-induced decision impacts among agents with strong task dependencies. It groups agents into adaptive local structures using density‑based clustering, aggregates intra‑group representations to smooth deviations, and then applies inter‑group attention to integrate information while respecting heterogeneous contributions. Experiments on noisy‑observation StarCraft II tasks confirm that SIGMA improves robustness to observation noise without sacrificing performance in clean environments.

By Li Mingqian