arXiv AI By Nikolaos Al. Papadopoulos, Konstantinos E. Psannis

Non-Stationarity Breaks Permutation Surrogates in Multi-Agent Reinforcement Learning: Diagnosis and Remedies

Read the original on arXiv AI →

The paper evaluates information‑theoretic measures for detecting directed influence in multi‑agent reinforcement learning by testing a guardrail in two games—a social dilemma and a coordination race—over 100 seeds. It shows that omitting the non‑stationary training transient leads to near‑perfect false‑positive rates, while simply excluding the transient is insufficient. The authors propose a block‑wise permutation null model that maintains the data structure and achieves low false‑positive rates (≈5%) while remaining highly sensitive to injected links.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Machine Learning
Aug 28

Shared Actors Need Not Share Critics: Effects of Value Mismatch in Parallel Reinforcement Learning

The paper investigates the problem of sharing a single critic across multiple parallel environments in reinforcement learning. It shows that when environments assign different expected returns to the same state, a shared critic must reconcile conflicting value targets, which can distort advantage estimates and misguide policy updates. The authors propose a simple fix—providing the critic with the environment index—demonstrating through bandit models and experiments on CartPole, MuJoCo, BipedalWalker, and 16 Procgen games that this conditional critic stabilizes learning and boosts returns, achieving a 40.8% improvement in aggregate normalized return on unseen levels.

By Zhenya Liu, Yang Meng, Zhuokai Zhao, Xuefeng Liu, Yuxin Chen