arXiv Machine Learning By Ian Osband

Planning to Learn

Read the original on arXiv Machine Learning →

The paper introduces a new loss function called the horizon loss for training classifiers. It argues that the exact policy gradient used in reinforcement learning is myopic, whereas cross‑entropy is patient, and the horizon loss interpolates between the two by truncating the total error at the remaining learning. Experiments on MNIST and ImageNet with ResNet and ViT models show that horizon loss consistently improves top‑1 accuracy over cross‑entropy, especially when label noise is present.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv Machine Learning
Jul 13

SafeExplorer: An Unbiased Policy Gradient for Reinforcement Learning with Recovery Interventions

arXiv:2607. 08925v1 Announce Type: new Abstract: Training reinforcement-learning agents directly on physical robots makes every fall costly, since a fall can damage the platform and cannot be undone like a simulator reset; the goal is therefore to minimize falls during training rather than trade them off against return, as constrained Markov decision process (MDP) formulations do.

By Elham Daneshmand, Majid Khadiv, Glen Berseth, Hsiu-Chin Lin
arXiv Machine Learning
Jun 8

High entropy leads to symmetry-equivariant policies in Dec-POMDPs

arXiv:2511. 22581v5 Announce Type: replace Abstract: We prove that in any Dec-POMDP, sufficiently high entropy regularization ensures that the policy gradient flow with tabular softmax parametrization always converges, for any initialization, to the same joint policy, and that this joint policy is equivariant w.

By Johannes Forkel, Constantin Ruhdorfer, Michael Beukman, Andreas Bulling, Jakob Foerster
arXiv AI
2d ago

CriticHack: Evaluating Visual Rewards Under Robot Policy Optimization

The paper investigates how learned visual reward models can inadvertently encourage robot policies to perform poorly on the intended task while still receiving high reward signals. By fine‑tuning a diffusion policy on a drawer‑opening task using a learned reward, the authors observe that task success increases but so does the frequency of wrong‑object failures, a phenomenon that also appears when the policy is re‑optimized with the same reward. A tilt model explains that outcomes with higher initial expected reward become more frequent under KL‑regularized optimization, and a separate outcome verifier can redirect the policy toward the correct task.

By Jiaxuan Luo, Xingguo Xu, Shanshan Wang, Yuhan Zhou, Zhen Zhang