arXiv AI By Tianle Zhong, Neiwen Ling, Yifan Pi, Zijun Wei, Tianshu Yu, Geoffrey Fox, Peng Wu, Xiao Yu

Diagnosing Training Inference Mismatch in LLM Reinforcement Learning via a Zero-Mismatch Reference

Read the original on arXiv AI →

The paper investigates Training‑Inference Mismatch (TIM) in large‑language‑model reinforcement learning, where rollout generation and policy optimization produce differing token probabilities despite identical model weights. By creating a zero‑mismatch diagnostic setting called VeXact, the authors isolate TIM and demonstrate that even minor token‑level numerical disagreements can trigger training collapse. They further show that TIM alters the effective optimization problem and propose remedies to mitigate its impact.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Machine Learning
Jun 30

The Mirage of Optimizing Training Policies: Monotonic Inference Policies as the Real Objective for LLM Reinforcement Learning

arXiv:2606. 29526v1 Announce Type: new Abstract: Reinforcement learning (RL) has gained growing attention in large language model (LLM) post-training, yet RL training remains fragile and can suffer from instability or collapse.

By Jing Liang, Hongyao Tang, Yi Ma, Yancheng He, Weixun Wang, Xiaoyang Li, Ju Huang, Wenbo Su, Jinyi Liu, Yan Zheng, Jianye Hao, Bo Zheng
arXiv Machine Learning
Sep 10

One Step, One Lead: Mitigating Higher-Order Interference in Multi-Domain Reinforcement Learning via Cross-Step Control

The paper introduces OSOL, a method for mitigating higher‑order interference in multi‑domain reinforcement learning. OSOL selects a focus domain each iteration, uses token‑level footprints from the previous checkpoint to rank rebound risk, and applies an adaptively scaled correction to the GRPO update. Experiments on Qwen3‑30B‑A3B show a 5.7% improvement over the best baseline without higher‑order differentiation.

By Zihan Lin, Xiaohan Wang, Jie Cao, Jiajun Chai, Guojun Yin, Wei Lin, Ran He