arXiv:2607. 21273v1 Announce Type: new Abstract: Dense per-step supervision is an appealing remedy for sparse-reward, long-horizon LLM agents: reward the agent for predicting its next observation, and memory should follow.
By Yu Wang
arXiv:2606. 18963v1 Announce Type: new Abstract: We study online reward-punishment learning when the environment provides no scalar reward or evaluative label.
By Zirong Li
The paper argues that in multi‑turn agentic reinforcement learning, credit assignment should be viewed as a coverage problem rather than a targeting problem. It introduces verifier information density (V_d) as a structural metric, showing that terminal‑state verifiers operate in a low‑V_d regime where targeting fails. Experiments on tau^2‑bench, BFCL, and ToolACE‑2‑8B demonstrate that uniformly distributing reward across all turns outperforms sparse, targeted rewards, and that full chain coverage is necessary for optimal performance.
By Chenyu Zhou, Qiliang Jiang, Shuning Wu, Xu Zhou
arXiv:2607. 05904v1 Announce Type: new Abstract: Training a language model against its own reference-free judgments (the premise of self-rewarding, self-play, and LLM-as-a-judge pipelines) assumes a model's verdict on a shown answer tracks correctness.
By Chenyu Zhou
The paper investigates selective on‑policy distillation, where a student model is trained only on token positions chosen by a selector. It demonstrates that the commonly used shared learning rate is not neutral: performance varies significantly with the learning rate for different selectors, leading to inconsistent comparisons. The authors attribute this selector‑rate entanglement to the selection process itself and recommend reporting the full arm‑by‑rate matrix for fair evaluation.
By Chencheng Zhu
arXiv:2605.11467v2 Announce Type: replace-cross
Abstract: Reasoning models post-hoc rationalize answers they have already committed to internally, producing chains of *reasoning theater*: deliberativ...
By Swapnil Parekh, Naman Goyal