arXiv AI

Interoceptive Attention as Dynamic Homeostatic Prioritization in a Foraging Agent

arXiv:2608. 04232v1 Announce Type: new Abstract: Biological systems must regulate competing needs under limited perceptual bandwidth, where sharpening one estimate costs the capacity to sharpen the others.

arXiv AI
Aug 18

When to Communicate: Belief Distributions and KL Divergence for Principled Gating in Multi-Agent RL

The paper proposes a principled communication strategy for multi‑agent reinforcement learning that gates messages based on the KL divergence between agents’ belief distributions over a latent world state. Each agent maintains a softmax belief derived from its LSTM hidden state and only communicates when disagreement exceeds a fixed threshold. Experiments on Predator‑Prey and MPE simple_spread show that this KL‑belief gating can match or surpass existing methods, improving performance and reducing variance in certain settings.

By Teoman Kaman
arXiv Machine Learning
Jul 27

Embodiment-Induced Coordination Regimes in Tabular Multi-Agent Q-Learning

arXiv:2601. 17454v2 Announce Type: replace-cross Abstract: Centralized value learning underlies a broad class of multi-agent reinforcement learning methods, but its claimed advantage is typically evaluated in settings that confound coordination structure with function approximation and partial observability.

By Muhammad Ahmed Atif, Nehal Naeem Haji, Mohammad Shahid Shaikh, Muhammad Ebad Atif
arXiv AI
Aug 25

Reinforcing the World's Edge: A Continual Learning Problem in the Multi-Agent-World Boundary

The paper studies a stationary decentralized Markov game where a focal agent experiences drifting rewards and dynamics due to learning peers, framing this as an agent‑centric continual reinforcement‑learning problem. It introduces the concept of an invariant core—maximal abstract patterns common to many successful trajectories—and proves a worst‑case conditioning theorem linking trajectory‑law drift to success coverage. The authors provide theoretical guarantees for survival horizon, first‑exit law, and regret, and validate their predictions with solvable models and empirical studies in continual control, cue‑MNIST, and Level‑Based Foraging.

By Dane Malenfant
arXiv Machine Learning
Aug 28

Shared Actors Need Not Share Critics: Effects of Value Mismatch in Parallel Reinforcement Learning

The paper investigates the problem of sharing a single critic across multiple parallel environments in reinforcement learning. It shows that when environments assign different expected returns to the same state, a shared critic must reconcile conflicting value targets, which can distort advantage estimates and misguide policy updates. The authors propose a simple fix—providing the critic with the environment index—demonstrating through bandit models and experiments on CartPole, MuJoCo, BipedalWalker, and 16 Procgen games that this conditional critic stabilizes learning and boosts returns, achieving a 40.8% improvement in aggregate normalized return on unseen levels.

By Zhenya Liu, Yang Meng, Zhuokai Zhao, Xuefeng Liu, Yuxin Chen
arXiv AI
Sep 4

Imagine-then-Plan: Agent Learning from Adaptive Lookahead with World Models

The paper introduces Imagine-then-Plan (ITP), a framework that lets agents learn by interacting with a learned world model to generate multi-step imagined trajectories. ITP features an adaptive lookahead mechanism that balances ultimate goals with task progress, producing richer signals about future outcomes. Experiments on various benchmarks show that ITP outperforms existing baselines, and analyses suggest the adaptive lookahead improves reasoning for complex tasks.

By Youwei Liu, Jian Wang, Hanlin Wang, Beichen Guo, Wenjie Li
arXiv AI
Jul 7

Regime-Conditional Stabilisation of LLM-Augmented Cooperative Multi-Agent Reinforcement Learning

arXiv:2607. 04470v1 Announce Type: cross Abstract: Large Language Models (LLMs) offer a natural interface for translating human objectives into reward signals for cooperative multi-agent reinforcement learning (MARL), yet the training-time dynamics of this integration remain poorly understood.

By Faid Keddouri, Sohaib Houhou, Aissa Boulmerka, Nadir Farhi