arXiv Machine Learning By Idil G\"ozel (University College London)

Learning Suffers More Than the Policy Class Under Partial Observability: A Closed-Form Analysis

Read the original on arXiv Machine Learning →

arXiv:2608. 07228v1 Announce Type: new Abstract: When a reinforcement learning agent cannot observe the full state, we usually blame its policies: it cannot see enough to represent a good one.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv Machine Learning
Jul 7

A Hierarchy of Policy Learning Problems

arXiv:2607. 03385v1 Announce Type: cross Abstract: Policy learning has received substantial attention with the goal of learning policies from observational data for decision-making.

By Hamsa Bastani, Osbert Bastani, Shihan Chen
arXiv Machine Learning
Aug 28

Shared Actors Need Not Share Critics: Effects of Value Mismatch in Parallel Reinforcement Learning

The paper investigates the problem of sharing a single critic across multiple parallel environments in reinforcement learning. It shows that when environments assign different expected returns to the same state, a shared critic must reconcile conflicting value targets, which can distort advantage estimates and misguide policy updates. The authors propose a simple fix—providing the critic with the environment index—demonstrating through bandit models and experiments on CartPole, MuJoCo, BipedalWalker, and 16 Procgen games that this conditional critic stabilizes learning and boosts returns, achieving a 40.8% improvement in aggregate normalized return on unseen levels.

By Zhenya Liu, Yang Meng, Zhuokai Zhao, Xuefeng Liu, Yuxin Chen
arXiv AI
Sep 25

Policy Complexity, Reaction Time, and Bounded Rationality in Reinforcement Learning

The paper introduces MI‑SARSA, an on‑policy temporal‑difference algorithm that incorporates mutual‑information regularization to model bounded rationality in reinforcement learning. By penalizing state‑specific deviations from a learned marginal action prior, the algorithm selectively uses state information only when the expected return outweighs the informational cost, yielding a reward‑complexity tradeoff. MI‑SARSA also predicts reaction times, showing that stronger information penalties lead to simpler policies, lower control costs, and faster responses, while regularization mitigates performance loss after environmental shifts at the expense of asymptotic return.

By James Wu, Chris R. Sims
arXiv AI
Sep 1

BCPPO: Bachelier-Inspired Constrained Proximal Policy Optimization for Tail-Risk-Aware Safe Reinforcement Learning

BCPPO is a new variant of Proximal Policy Optimization that uses Bachelier-inspired cost‑prediction networks to generate a smooth penalty based on disagreement among critics. The method keeps temporal‑difference learning unchanged, applies a saturation‑aware controller to manage cost penalties, and deploys only the policy network. Across extensive experiments, BCPPO outperforms comparators in achieving higher mean returns while maintaining lower or comparable CVaR in all tested tasks.

By Dongsheng Hou, Yanqiao Chen, Yuhan Rui