arXiv Machine Learning By Jim Allchin

When Does Reward Teach State? A Hidden-Automaton Instrument and the Group-Language Boundary

Read the original on arXiv Machine Learning →

arXiv:2607. 11953v1 Announce Type: new Abstract: Does a reinforcement-learning agent that earns high reward represent its task's latent state, or only a reward-correlated shortcut?

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv AI
Sep 2

Explore More, Drift Less: Outcome-Only Reinforcement Learning Can Suffice for Long-Horizon Interactive Agents

The paper proposes CANOPY, a minimalist reinforcement learning protocol that addresses two common pitfalls—signal starvation and policy drift—in outcome‑only RL for long‑horizon interactive tasks. By scaling same‑task exploration, keeping updates on‑policy, and anchoring updates with KL divergence, CANOPY enables a Qwen3‑14B agent to achieve top leaderboard results on the AppWorld coding benchmark without auxiliary supervision or elaborate scaffolding. The approach also improves performance on SWE‑bench for a Qwen3.5‑9B model.

By Liming Pu, Xiaoxia Li, Yifu Liu, Teng Cao, Bin Yang
arXiv AI
3d ago

From Imitation to Reward Discovery: On-Policy Warmup for Agentic RL

The paper introduces On‑Policy Warmup (OPW), a teacher‑guided training stage where a student agent learns from a teacher on its own interaction trajectories before switching to reinforcement learning with verifiable rewards (RLVR). OPW differs from traditional imitation by focusing on states generated by the student’s own decisions, including imperfect actions and recovery situations. The authors provide a theoretical link between on‑policy reverse‑KL distillation and trajectory‑level distribution matching, showing that, under a competent teacher and low distillation loss, OPW can lower bound initial verifier success and reduce reward‑discovery complexity, thereby accelerating RLVR performance.

By Yitong Qiao, Tiantian He, Lei Liu, Yue Shen, Jian Wang, Jinjie Gu, Zhixuan Chu
arXiv Computation and Language
Aug 28

Reward-Informed Sparse Autoencoders and the Solution-Completeness Confound

The paper introduces Reward‑Informed Sparse Autoencoders (RI‑SAEs), which use reinforcement‑learning rewards to curate data for training sparse autoencoders on language‑model activations. On Llama‑3.1‑8B, a sparse subset of features separates high‑reward from low‑reward reasoning continuations, but control experiments show this separation largely reflects solution completeness rather than true reasoning quality. The authors conclude that reward filtering can cheaply reuse RL signals for interpretability, though most of the discovered features capture completion form rather than deep reasoning.

By Tanvi Nagilla, Alexander Jameson, Daniel Manta, Shayaan Uddin