The paper proposes a principled communication strategy for multi‑agent reinforcement learning that gates messages based on the KL divergence between agents’ belief distributions over a latent world state. Each agent maintains a softmax belief derived from its LSTM hidden state and only communicates when disagreement exceeds a fixed threshold. Experiments on Predator‑Prey and MPE simple_spread show that this KL‑belief gating can match or surpass existing methods, improving performance and reducing variance in certain settings.
By Teoman Kaman
arXiv:2609.34373v2 Announce Type: replace
Abstract: Rate-limited multi-agent teams raise three questions the emergent-communication literature has answered only empirically: what an optimal message s...
By Mihir Chauhan, Aniket Bera
arXiv:2601. 17454v2 Announce Type: replace-cross Abstract: Centralized value learning underlies a broad class of multi-agent reinforcement learning methods, but its claimed advantage is typically evaluated in settings that confound coordination structure with function approximation and partial observability.
By Muhammad Ahmed Atif, Nehal Naeem Haji, Mohammad Shahid Shaikh, Muhammad Ebad Atif
arXiv:2609.39342v1 Announce Type: cross
Abstract: Intrinsic motivation plays a central role in adaptive and goal-directed behavior by conferring agents reward-independent objectives and biases useful...
By Manolis Mylonas, Rub\'en Moreno Bote
Intrinsic motivation plays a central role in adaptive and goal-directed behavior by conferring agents reward-independent objectives and biases useful to act in noisy and uncertain environments. Active...
The paper studies a stationary decentralized Markov game where a focal agent experiences drifting rewards and dynamics due to learning peers, framing this as an agent‑centric continual reinforcement‑learning problem. It introduces the concept of an invariant core—maximal abstract patterns common to many successful trajectories—and proves a worst‑case conditioning theorem linking trajectory‑law drift to success coverage. The authors provide theoretical guarantees for survival horizon, first‑exit law, and regret, and validate their predictions with solvable models and empirical studies in continual control, cue‑MNIST, and Level‑Based Foraging.
By Dane Malenfant
The paper investigates the problem of sharing a single critic across multiple parallel environments in reinforcement learning. It shows that when environments assign different expected returns to the same state, a shared critic must reconcile conflicting value targets, which can distort advantage estimates and misguide policy updates. The authors propose a simple fix—providing the critic with the environment index—demonstrating through bandit models and experiments on CartPole, MuJoCo, BipedalWalker, and 16 Procgen games that this conditional critic stabilizes learning and boosts returns, achieving a 40.8% improvement in aggregate normalized return on unseen levels.
By Zhenya Liu, Yang Meng, Zhuokai Zhao, Xuefeng Liu, Yuxin Chen
arXiv:2608. 11506v1 Announce Type: cross Abstract: Adaptive behavior under partial observability depends on internal organization that carries information beyond the current observation.
By Frederick Hayes III
arXiv:2607. 06503v1 Announce Type: new Abstract: Large language model (LLM) agents solving multi-step tasks frequently commit to trajectories that are doomed to fail, yet continue to consume substantial inference compute before the failure becomes observable.
By Kai Ruan, Zihe Huang, Ziqi Zhou, Qianshan Wei, Xuan Wang, Hao Sun
The paper introduces Imagine-then-Plan (ITP), a framework that lets agents learn by interacting with a learned world model to generate multi-step imagined trajectories. ITP features an adaptive lookahead mechanism that balances ultimate goals with task progress, producing richer signals about future outcomes. Experiments on various benchmarks show that ITP outperforms existing baselines, and analyses suggest the adaptive lookahead improves reasoning for complex tasks.
By Youwei Liu, Jian Wang, Hanlin Wang, Beichen Guo, Wenjie Li
arXiv:2608.28046v1 Announce Type: cross
Abstract: Collective behaviour in living systems is usually modelled as the outcome of a \emph{direct} social drive: agents are rewarded, or hard-wired, to ali...
By Gorka Mu\~noz-Gil, Andrea L\'opez-Incera, Vide Ramsten, Giovanni Volpe, Thomas M\"uller, Hans J. Briegel
arXiv:2607. 04470v1 Announce Type: cross Abstract: Large Language Models (LLMs) offer a natural interface for translating human objectives into reward signals for cooperative multi-agent reinforcement learning (MARL), yet the training-time dynamics of this integration remain poorly understood.
By Faid Keddouri, Sohaib Houhou, Aissa Boulmerka, Nadir Farhi