arXiv Machine Learning By Zhipeng Zhang

Stable but Wrong: When Learning Stabilizes Away from the Truth

Read the original on arXiv Machine Learning →

The paper introduces the concept of Stable but Wrong (SBW), describing situations where a learning process appears stable yet converges to a solution that is systematically biased away from a true objective. Using a minimal strongly convex model, the authors demonstrate that persistent bias in update directions can shift the convergence point from the optimal solution. Experiments across reinforcement learning, supervised learning, and continual fine-tuning of large language models reveal a recurring disconnect between normal optimization behavior and correctness under both static and feedback‑coupled biases, and show that recovery interventions can still modify subsequent learning trajectories.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv Machine Learning
Jun 30

The Mirage of Optimizing Training Policies: Monotonic Inference Policies as the Real Objective for LLM Reinforcement Learning

arXiv:2606. 29526v1 Announce Type: new Abstract: Reinforcement learning (RL) has gained growing attention in large language model (LLM) post-training, yet RL training remains fragile and can suffer from instability or collapse.

By Jing Liang, Hongyao Tang, Yi Ma, Yancheng He, Weixun Wang, Xiaoyang Li, Ju Huang, Wenbo Su, Jinyi Liu, Yan Zheng, Jianye Hao, Bo Zheng
arXiv Machine Learning
Jun 5

Extreme Region Policy Distillation

arXiv:2605. 25582v2 Announce Type: replace Abstract: Reinforcement learning for large language models faces a fundamental trade-off between sample efficiency and asymptotic performance: strictly on-policy methods discard trajectories after a single update, while off-policy reuse introduces distribution mismatch that existing trust-region techniques mitigate primarily by enforcing conservative optimization, often leaving rich training signals underutilized.

By Changyu Chen, Xiting Wang, Rui Yan