arXiv Machine Learning
Aug 4

Toward Plasticity-Preserving KL Regularization for Capability Retention in LLM Reinforcement Learning

arXiv:2608. 01743v1 Announce Type: new Abstract: Reinforcement learning (RL) has become a central paradigm for large language model (LLM) post-training, but optimization toward new objectives can degrade capabilities already present in the base model.

By Li Wang, Xiaodong Lu, Xiaohan Wang, Jiajun Chai, Wei Lin, Tianhao Peng, Guojun Yin
arXiv Machine Learning
1d ago

Poincar\'e Meets Bellman: Revisable Memory, Operational Quotients, and Evidence-Supported Learning in Changing Environments

The paper presents a finite‑model synthesis that integrates operational state abstraction with optimal control within a stability‑evidence‑revision (SER) framework, termed "Poincaré meets Bellman." It distinguishes qualitative dynamics for reusable action‑response structures from dynamic programming that governs acquisition, retention, reuse, merging, and forgetting, and introduces a Bellman recursion over the joint law of hidden state and deployed memory. The authors derive explicit retention rules, demonstrate how factor sharing and informative observations improve identification, and verify coding and retention calculations through finite enumerations.

By Xin Li
arXiv AI
Aug 25

Functional compatibility as a determinant of persistent neural learning

The paper demonstrates that functional compatibility—how well new learning can coexist with behavior that must be preserved—is a causal determinant of persistent neural learning. By deliberately altering compatibility between matched neural states, the authors show that persistent learning can be measured under a common retention requirement across different learning directions, architectures, and data modalities. The study reveals that learning rules vary in how efficiently they exploit compatibility, and that retention constraints and finite updates limit what can be stored, with nonlinear geometry ultimately preventing full compatibility realization.

By Hossein Javidnia
arXiv AI
Aug 26

Revelation Control

Revelation Control studies how to price interventions that reveal hidden state only when the revealed distinctions can alter a consequential decision, while separately accounting for any useful progress the intervention itself creates. The authors develop a framework for learning systems that defines decision‑sufficient revelation, revelation depth, and a cost‑adjusted factorization criterion, and they provide a target‑independent protocol for model‑specific instantiation. Experiments on Qwen2.5‑7B and Mistral‑7B‑v0.3 show that deeper future‑learning probes have positive decision value and that productive reuse yields strict equal‑compute utility advantages, supporting a structural transfer of the decision theory and evaluation protocol across architectures.

By Qinyou Wang