arXiv AI
3d ago

Diagnosing Training Inference Mismatch in LLM Reinforcement Learning via a Zero-Mismatch Reference

The paper investigates Training‑Inference Mismatch (TIM) in large‑language‑model reinforcement learning, where rollout generation and policy optimization produce differing token probabilities despite identical model weights. By creating a zero‑mismatch diagnostic setting called VeXact, the authors isolate TIM and demonstrate that even minor token‑level numerical disagreements can trigger training collapse. They further show that TIM alters the effective optimization problem and propose remedies to mitigate its impact.

By Tianle Zhong, Neiwen Ling, Yifan Pi, Zijun Wei, Tianshu Yu, Geoffrey Fox, Peng Wu, Xiao Yu
arXiv AI
Sep 17

REVERSAL-BENCH: A Reversibility Axis and Reset Oracle for Measuring the Reset-Free RL Cliff

REVERSAL-BENCH is a benchmark that introduces a continuous reversibility parameter ρ∈[0,1] and a reset oracle to evaluate how well reinforcement learning agents can recover from irreversible states across eight manipulation tasks in five physics engines. Experiments show a sharp reversibility cliff: reset‑free agents become trapped in irrecoverable states as ρ increases, while episodic agents continue learning steadily. The benchmark also provides a large multi‑simulator dataset and demonstrates that safety shields can predict recoverability but only succeed when the agent can avoid the trap.

By Riyaaz Shaik, Chandru Venkataraman
arXiv Machine Learning
Sep 24

Repairability of Inexact Solvers in Recursive State Estimation with Machine Learning

The paper investigates how approximate numerical solvers used in recursive state estimation can be repaired using bounded corrections, characterizing when such corrections meet local admissibility tolerances and how they influence finite‑horizon covariance. It derives error identities that separate solve error from gain drift, revealing quartic and sixth‑order contributions to the covariance response. The framework is applied to a power‑grid tolerance study, showing that learned corrections reduce the required conjugate‑gradient iterations, and it demonstrates a unified interface for classical, quantum, and hybrid solvers.

By Yanjun Ji, Dennis Willsch, Orkun \c{S}ensebat, Priyanka Arkalgud Ganeshamurthy, Zhi Pei, M. Sahnawaz Alam, Ivelina Stoyanova, Frank K. Wilhelm, Bo Zhao, Chao Wang, Kristel Michielsen
Hugging Face Trending Papers
Jul 13

The Equilibrium Is the Initialization: Lazy Identity Collapse in Physics-Structured Deep Equilibrium Reasoning

Deep equilibrium models promise input-adaptive implicit computation: harder problems should demand more solver iterations, and the solved equilibrium should encode the result of genuine iterative inference. We report a cautionary study of a port-Hamiltonian DEQ with a learned initialization on two reasoning tasks -- ProofWriter entailment over frozen DeBERTa embeddings and a BFS-verified graph-reachability benchmark -- in which the implicit computation is a silent no-op.