arXiv:2608. 01130v1 Announce Type: new Abstract: A broad range of models face the mismatch where they are updated through trajectory losses but are evaluated by downstream task reward.
By Yuyang Shen
arXiv:2607. 08925v1 Announce Type: new Abstract: Training reinforcement-learning agents directly on physical robots makes every fall costly, since a fall can damage the platform and cannot be undone like a simulator reset; the goal is therefore to minimize falls during training rather than trade them off against return, as constrained Markov decision process (MDP) formulations do.
By Elham Daneshmand, Majid Khadiv, Glen Berseth, Hsiu-Chin Lin
arXiv:2511. 22581v5 Announce Type: replace Abstract: We prove that in any Dec-POMDP, sufficiently high entropy regularization ensures that the policy gradient flow with tabular softmax parametrization always converges, for any initialization, to the same joint policy, and that this joint policy is equivariant w.
By Johannes Forkel, Constantin Ruhdorfer, Michael Beukman, Andreas Bulling, Jakob Foerster
arXiv:2607. 01490v1 Announce Type: cross Abstract: Reinforcement learning post-training dramatically improves LLM reasoning, but suffers from training instability and diversity collapse.
By Juliette Decugis, Sean O'Brien, Francis Bach, Gabriel Synnaeve, Taco Cohen
arXiv:2606. 18963v1 Announce Type: new Abstract: We study online reward-punishment learning when the environment provides no scalar reward or evaluative label.
By Zirong Li
The paper investigates how learned visual reward models can inadvertently encourage robot policies to perform poorly on the intended task while still receiving high reward signals. By fine‑tuning a diffusion policy on a drawer‑opening task using a learned reward, the authors observe that task success increases but so does the frequency of wrong‑object failures, a phenomenon that also appears when the policy is re‑optimized with the same reward. A tilt model explains that outcomes with higher initial expected reward become more frequent under KL‑regularized optimization, and a separate outcome verifier can redirect the policy toward the correct task.
By Jiaxuan Luo, Xingguo Xu, Shanshan Wang, Yuhan Zhou, Zhen Zhang