arXiv Machine Learning

Planning to Learn

The paper introduces a new loss function called the horizon loss for training classifiers. It argues that the exact policy gradient used in reinforcement learning is myopic, whereas cross‑entropy is patient, and the horizon loss interpolates between the two by truncating the total error at the remaining learning. Experiments on MNIST and ImageNet with ResNet and ViT models show that horizon loss consistently improves top‑1 accuracy over cross‑entropy, especially when label noise is present.

arXiv Machine Learning
Jul 13

SafeExplorer: An Unbiased Policy Gradient for Reinforcement Learning with Recovery Interventions

arXiv:2607. 08925v1 Announce Type: new Abstract: Training reinforcement-learning agents directly on physical robots makes every fall costly, since a fall can damage the platform and cannot be undone like a simulator reset; the goal is therefore to minimize falls during training rather than trade them off against return, as constrained Markov decision process (MDP) formulations do.

By Elham Daneshmand, Majid Khadiv, Glen Berseth, Hsiu-Chin Lin
arXiv Machine Learning
Jun 8

High entropy leads to symmetry-equivariant policies in Dec-POMDPs

arXiv:2511. 22581v5 Announce Type: replace Abstract: We prove that in any Dec-POMDP, sufficiently high entropy regularization ensures that the policy gradient flow with tabular softmax parametrization always converges, for any initialization, to the same joint policy, and that this joint policy is equivariant w.

By Johannes Forkel, Constantin Ruhdorfer, Michael Beukman, Andreas Bulling, Jakob Foerster
arXiv AI
2d ago

CriticHack: Evaluating Visual Rewards Under Robot Policy Optimization

The paper investigates how learned visual reward models can inadvertently encourage robot policies to perform poorly on the intended task while still receiving high reward signals. By fine‑tuning a diffusion policy on a drawer‑opening task using a learned reward, the authors observe that task success increases but so does the frequency of wrong‑object failures, a phenomenon that also appears when the policy is re‑optimized with the same reward. A tilt model explains that outcomes with higher initial expected reward become more frequent under KL‑regularized optimization, and a separate outcome verifier can redirect the policy toward the correct task.

By Jiaxuan Luo, Xingguo Xu, Shanshan Wang, Yuhan Zhou, Zhen Zhang
arXiv AI
Sep 1

BCPPO: Bachelier-Inspired Constrained Proximal Policy Optimization for Tail-Risk-Aware Safe Reinforcement Learning

BCPPO is a new variant of Proximal Policy Optimization that uses Bachelier-inspired cost‑prediction networks to generate a smooth penalty based on disagreement among critics. The method keeps temporal‑difference learning unchanged, applies a saturation‑aware controller to manage cost penalties, and deploys only the policy network. Across extensive experiments, BCPPO outperforms comparators in achieving higher mean returns while maintaining lower or comparable CVaR in all tested tasks.

By Dongsheng Hou, Yanqiao Chen, Yuhan Rui
arXiv AI
Aug 3

Deconstructing Off-Policy Ratios: Entropy-Scaled Trust Regions for Asynchronous Reinforcement Learning

arXiv:2607. 22186v2 Announce Type: replace Abstract: Asynchronous reinforcement learning (RL) accelerates large language model (LLM) post-training by overlapping rollout generation with policy optimization, but the resulting stale, off-policy data can destabilize optimization and ultimately cause policy collapse.

By Guanqun Zhao, Zijun Xie, Binbin Zheng, Enlei Gong, Jiafeng Lu, Yehan Yang, Aoqi Hu, Zeyu Chen
arXiv Machine Learning
5d ago

Backward-State Policy Is Part of the Learning Algorithm

The paper argues that the policy governing how tensors are rounded and reused during the backward pass—termed the backward‑state policy—is an integral part of the learning algorithm, not merely a memory or precision detail. Experiments with 390 M‑parameter models show that whether the backward pass reuses a forward’s rounded output or generates a new rounding can decisively affect training success, even when copy accuracy is high. The authors propose a method to determine, for each operator, which value should be read or substituted to preserve the correct gradient, and validate these predictions on PyTorch and Transformer Engine.

By Shuxiao Xie, Shuyang Xie, Dezhi Ran, Wei Yang, Tao Xie