arXiv Machine Learning By Xujun Che, Yuchen Yuan, Weida Zhao, Chenyang Yu

Abstention as an Action Can Kill Both the Reward Gradient and the KL Anchor: Collapse Law and Repair for Error-Penalized Reinforcement Learning

Read the original on arXiv Machine Learning →

arXiv:2608. 00301v1 Announce Type: new Abstract: Error-penalized scoring rules ($+1$ for a correct answer, $-\lambda$ for a wrong one, $0$ for abstaining) are increasingly prescribed against hallucination: a rational agent facing such a rule answers exactly when its correctness probability exceeds Chow's threshold $t^\ast=\lambda/(1+\lambda)$.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv AI
Jun 30

Pessimism's Paradox: Conservative Offline Training Amplifies Reward Hacking During Online Adaptation in Reasoning Models

arXiv:2606. 30627v1 Announce Type: cross Abstract: Conservative offline training is widely advocated as a safe foundation for subsequent online adaptation: if a policy stays close to well-supported behaviour, the argument goes, it is less likely to exploit imperfections in a learned reward model.

By Subramanyam Sahoo, Aman Chadha, Vinija Jain, Divya Chaudhary
arXiv AI
Aug 18

When to Communicate: Belief Distributions and KL Divergence for Principled Gating in Multi-Agent RL

The paper proposes a principled communication strategy for multi‑agent reinforcement learning that gates messages based on the KL divergence between agents’ belief distributions over a latent world state. Each agent maintains a softmax belief derived from its LSTM hidden state and only communicates when disagreement exceeds a fixed threshold. Experiments on Predator‑Prey and MPE simple_spread show that this KL‑belief gating can match or surpass existing methods, improving performance and reducing variance in certain settings.

By Teoman Kaman