arXiv Machine Learning

High entropy leads to symmetry-equivariant policies in Dec-POMDPs

arXiv:2511. 22581v5 Announce Type: replace Abstract: We prove that in any Dec-POMDP, sufficiently high entropy regularization ensures that the policy gradient flow with tabular softmax parametrization always converges, for any initialization, to the same joint policy, and that this joint policy is equivariant w.

arXiv Machine Learning
Aug 28

Shared Actors Need Not Share Critics: Effects of Value Mismatch in Parallel Reinforcement Learning

The paper investigates the problem of sharing a single critic across multiple parallel environments in reinforcement learning. It shows that when environments assign different expected returns to the same state, a shared critic must reconcile conflicting value targets, which can distort advantage estimates and misguide policy updates. The authors propose a simple fix—providing the critic with the environment index—demonstrating through bandit models and experiments on CartPole, MuJoCo, BipedalWalker, and 16 Procgen games that this conditional critic stabilizes learning and boosts returns, achieving a 40.8% improvement in aggregate normalized return on unseen levels.

By Zhenya Liu, Yang Meng, Zhuokai Zhao, Xuefeng Liu, Yuxin Chen
arXiv AI
Aug 3

Deconstructing Off-Policy Ratios: Entropy-Scaled Trust Regions for Asynchronous Reinforcement Learning

arXiv:2607. 22186v2 Announce Type: replace Abstract: Asynchronous reinforcement learning (RL) accelerates large language model (LLM) post-training by overlapping rollout generation with policy optimization, but the resulting stale, off-policy data can destabilize optimization and ultimately cause policy collapse.

By Guanqun Zhao, Zijun Xie, Binbin Zheng, Enlei Gong, Jiafeng Lu, Yehan Yang, Aoqi Hu, Zeyu Chen
arXiv Machine Learning
Jul 13

SafeExplorer: An Unbiased Policy Gradient for Reinforcement Learning with Recovery Interventions

arXiv:2607. 08925v1 Announce Type: new Abstract: Training reinforcement-learning agents directly on physical robots makes every fall costly, since a fall can damage the platform and cannot be undone like a simulator reset; the goal is therefore to minimize falls during training rather than trade them off against return, as constrained Markov decision process (MDP) formulations do.

By Elham Daneshmand, Majid Khadiv, Glen Berseth, Hsiu-Chin Lin
arXiv AI
Jun 2

S-SPPO: Semantic-Calibrated Self-Play Preference Optimization

arXiv:2606. 01561v1 Announce Type: new Abstract: Aligning Large Language Models (LLMs) with human preferences is often formulated via Direct Preference Optimization (DPO).

By Xiwen Chen, Wenhui Zhu, Jingjing Wang, Peijie Qiu, Zhipeng Wang, Huayu Li, ZhengXiao He, Xuanzhao Dong, Prayag Tiwari, Mingkun Xu, Yujian Xiong, Feng Luo, Abolfazl Razi, Brendan Hogan Rappazzo, Anderson Schneider, Yuriy Nevmyvaka
arXiv AI
6d ago

PolicyAttention: Softmax Attention Implements Policy Mirror Descent for Closed-Loop Control

The paper investigates whether causal softmax attention can realize policy mirror descent (PMD) as a repeated controller rather than a one‑step algebraic identity. It constructs a fixed causal‑softmax actor–environment–one‑step‑critic protocol, detailing actor, routing, sampling, and normalization residuals, and shows that a frozen one‑step audit model closely approximates PMD. Empirical results demonstrate that the learned actor with an exact one‑step critic achieves median policy loss only about 5% higher than the exact PMD oracle across multiple control settings.

By Yuhe Sui, Yingzhi Tang, Shufang Chen
Hugging Face Trending Papers
Jun 28

Dead-Direction Conditioners: Gauge-Equivariant Preconditioning for Deep Networks

A deep network's loss is invariant to continuous symmetries of its parameters: the logit shift, the ReLU rescaling, the LayerNorm scale, the per-head attention rotation. Adam's per-coordinate preconditioner drifts along each symmetry orbit, which pulls the trajectory off the symmetry quotient where the optimization lives and blurs the singular-learning rate the quotient makes readable.

arXiv AI
Sep 1

PokaiTrainer: Scaling Belief-State Search to Competitive Pok\'emon VGC

PokaiTrainer is a competitive Pokémon VGC agent that scales belief‑state search to handle simultaneous, large joint action spaces and stochastic outcomes. The system uses PokaiEngine, a Rust battle engine that efficiently enumerates joint action outcomes with high accuracy, and adapts Student of Games to solve each decision as a Bayesian matrix game under a compute budget. In live Showdown play, the agent achieved a 59% win rate against a human field averaging ~1320 Elo, reaching an Elo band of 1350‑1400 and briefly entering the top 500 of the format.

By Max Yu