arXiv:2607. 11953v3 Announce Type: replace Abstract: Does a reinforcement-learning agent that earns high reward actually learn its task's hidden state, or only a shortcut that correlates with reward?
By James E. Allchin
arXiv:2608. 05111v1 Announce Type: new Abstract: In partially observable reinforcement learning, agents face a dual bottleneck: they must explore to encounter rewarding states and retain that experience in memory to optimize their policies.
By Jai Malegaonkar, Rohan Patil, Henrik I. Christensen
The paper proposes CANOPY, a minimalist reinforcement learning protocol that addresses two common pitfalls—signal starvation and policy drift—in outcome‑only RL for long‑horizon interactive tasks. By scaling same‑task exploration, keeping updates on‑policy, and anchoring updates with KL divergence, CANOPY enables a Qwen3‑14B agent to achieve top leaderboard results on the AppWorld coding benchmark without auxiliary supervision or elaborate scaffolding. The approach also improves performance on SWE‑bench for a Qwen3.5‑9B model.
By Liming Pu, Xiaoxia Li, Yifu Liu, Teng Cao, Bin Yang
The paper introduces On‑Policy Warmup (OPW), a teacher‑guided training stage where a student agent learns from a teacher on its own interaction trajectories before switching to reinforcement learning with verifiable rewards (RLVR). OPW differs from traditional imitation by focusing on states generated by the student’s own decisions, including imperfect actions and recovery situations. The authors provide a theoretical link between on‑policy reverse‑KL distillation and trajectory‑level distribution matching, showing that, under a competent teacher and low distillation loss, OPW can lower bound initial verifier success and reduce reward‑discovery complexity, thereby accelerating RLVR performance.
By Yitong Qiao, Tiantian He, Lei Liu, Yue Shen, Jian Wang, Jinjie Gu, Zhixuan Chu
The paper introduces Reward‑Informed Sparse Autoencoders (RI‑SAEs), which use reinforcement‑learning rewards to curate data for training sparse autoencoders on language‑model activations. On Llama‑3.1‑8B, a sparse subset of features separates high‑reward from low‑reward reasoning continuations, but control experiments show this separation largely reflects solution completeness rather than true reasoning quality. The authors conclude that reward filtering can cheaply reuse RL signals for interpretability, though most of the discovered features capture completion form rather than deep reasoning.
By Tanvi Nagilla, Alexander Jameson, Daniel Manta, Shayaan Uddin
arXiv:2607. 21273v1 Announce Type: new Abstract: Dense per-step supervision is an appealing remedy for sparse-reward, long-horizon LLM agents: reward the agent for predicting its next observation, and memory should follow.
By Yu Wang