arXiv:2607. 11953v1 Announce Type: new Abstract: Does a reinforcement-learning agent that earns high reward represent its task's latent state, or only a reward-correlated shortcut?
By Jim Allchin
arXiv:2607. 18966v1 Announce Type: new Abstract: Language models trained with reinforcement learning may learn to optimize the grader's judgment rather than the intended objective.
By Axel H{\o}jmark, J\'er\'emy Scheurer, Evgenia Nitishinskaya, Felix Hofst\"atter, Jason Wolfe, Theodore Ehrenborg, Bronson Schoen, Alexander Meinke
arXiv:2607. 05904v1 Announce Type: new Abstract: Training a language model against its own reference-free judgments (the premise of self-rewarding, self-play, and LLM-as-a-judge pipelines) assumes a model's verdict on a shown answer tracks correctness.
By Chenyu Zhou
arXiv:2608. 05111v1 Announce Type: new Abstract: In partially observable reinforcement learning, agents face a dual bottleneck: they must explore to encounter rewarding states and retain that experience in memory to optimize their policies.
By Jai Malegaonkar, Rohan Patil, Henrik I. Christensen
arXiv:2607. 21273v1 Announce Type: new Abstract: Dense per-step supervision is an appealing remedy for sparse-reward, long-horizon LLM agents: reward the agent for predicting its next observation, and memory should follow.
By Yu Wang
arXiv:2608. 01425v1 Announce Type: cross Abstract: Training LLM-based multi-agent systems with multi-agent reinforcement learning is rapidly gaining traction, and a parallel line of work argues that such systems should be judged by their behavior, not only their reward.
By Yi Mao, Andrew Perrault