arXiv:2607. 21273v2 Announce Type: replace Abstract: Dense per-step supervision is the standard remedy for sparse-reward long-horizon LLM agents: reward the policy for predicting its next observation, which looks provably safe under potential-based shaping.
By Yu Wang
arXiv:2609.00892v1 Announce Type: new
Abstract: Rubric-based reinforcement learning decomposes open-ended instructions into prompt-specific, flexible rubrics, making it better suited than reinforceme...
By Siyuan Li, Xinxin Song, Chen Ruinian, Jingjing Fan, Tingxiong Xiao, Yangen Hu, Ke Zeng, Jinli Suo
arXiv:2606. 30627v1 Announce Type: cross Abstract: Conservative offline training is widely advocated as a safe foundation for subsequent online adaptation: if a policy stays close to well-supported behaviour, the argument goes, it is less likely to exploit imperfections in a learned reward model.
By Subramanyam Sahoo, Aman Chadha, Vinija Jain, Divya Chaudhary
arXiv:2607. 07435v1 Announce Type: cross Abstract: Agents acting on our behalf in the real world (e.
By Bojie Li, Noah Shi
arXiv:2607. 21273v1 Announce Type: new Abstract: Dense per-step supervision is an appealing remedy for sparse-reward, long-horizon LLM agents: reward the agent for predicting its next observation, and memory should follow.
By Yu Wang
The paper proposes a principled communication strategy for multi‑agent reinforcement learning that gates messages based on the KL divergence between agents’ belief distributions over a latent world state. Each agent maintains a softmax belief derived from its LSTM hidden state and only communicates when disagreement exceeds a fixed threshold. Experiments on Predator‑Prey and MPE simple_spread show that this KL‑belief gating can match or surpass existing methods, improving performance and reducing variance in certain settings.
By Teoman Kaman