arXiv AI

Non-Stationarity Breaks Permutation Surrogates in Multi-Agent Reinforcement Learning: Diagnosis and Remedies

The paper evaluates information‑theoretic measures for detecting directed influence in multi‑agent reinforcement learning by testing a guardrail in two games—a social dilemma and a coordination race—over 100 seeds. It shows that omitting the non‑stationary training transient leads to near‑perfect false‑positive rates, while simply excluding the transient is insufficient. The authors propose a block‑wise permutation null model that maintains the data structure and achieves low false‑positive rates (≈5%) while remaining highly sensitive to injected links.

arXiv Machine Learning
Aug 28

Shared Actors Need Not Share Critics: Effects of Value Mismatch in Parallel Reinforcement Learning

The paper investigates the problem of sharing a single critic across multiple parallel environments in reinforcement learning. It shows that when environments assign different expected returns to the same state, a shared critic must reconcile conflicting value targets, which can distort advantage estimates and misguide policy updates. The authors propose a simple fix—providing the critic with the environment index—demonstrating through bandit models and experiments on CartPole, MuJoCo, BipedalWalker, and 16 Procgen games that this conditional critic stabilizes learning and boosts returns, achieving a 40.8% improvement in aggregate normalized return on unseen levels.

By Zhenya Liu, Yang Meng, Zhuokai Zhao, Xuefeng Liu, Yuxin Chen
arXiv Machine Learning
4d ago

Frozen Judges, Moving Agents: Version-Dependent LLM-Judge Error and the Limits of Judge-Assisted Agent Evaluation

The paper investigates how language‑model judges can make version‑dependent errors when evaluating upgraded agents. Using 35 public coding‑agent submissions, two customer‑service agents, and over a thousand expert‑labeled trajectories, the authors show that fixed judges often reject task‑conditioned error invariance and can incorrectly approve failed patches, especially as agent capability increases. Paired audits of current outputs reduce interval width only marginally, and the study concludes that independent human patch review is still necessary.

By Jiapeng Li
arXiv Machine Learning
Sep 22

A Shared Learning Rate Is Not a Neutral Control in Selective On-Policy Distillation

The paper investigates selective on‑policy distillation, where a student model is trained only on token positions chosen by a selector. It demonstrates that the commonly used shared learning rate is not neutral: performance varies significantly with the learning rate for different selectors, leading to inconsistent comparisons. The authors attribute this selector‑rate entanglement to the selection process itself and recommend reporting the full arm‑by‑rate matrix for fair evaluation.

By Chencheng Zhu