Diagnosing Faults in Reinforcement Learning Simulators and World Models with Canonical Polynomial Invariants
Read the original on arXiv AI →The Flow has not summarised this story yet — read it at arXiv AI.
The Flow has not summarised this story yet — read it at arXiv AI.
The paper investigates Training‑Inference Mismatch (TIM) in large‑language‑model reinforcement learning, where rollout generation and policy optimization produce differing token probabilities despite identical model weights. By creating a zero‑mismatch diagnostic setting called VeXact, the authors isolate TIM and demonstrate that even minor token‑level numerical disagreements can trigger training collapse. They further show that TIM alters the effective optimization problem and propose remedies to mitigate its impact.
REVERSAL-BENCH is a benchmark that introduces a continuous reversibility parameter ρ∈[0,1] and a reset oracle to evaluate how well reinforcement learning agents can recover from irreversible states across eight manipulation tasks in five physics engines. Experiments show a sharp reversibility cliff: reset‑free agents become trapped in irrecoverable states as ρ increases, while episodic agents continue learning steadily. The benchmark also provides a large multi‑simulator dataset and demonstrates that safety shields can predict recoverability but only succeed when the agent can avoid the trap.
arXiv:2608. 05702v1 Announce Type: new Abstract: Scientific machine learning commonly validates models at the level of a subdomain, a benchmark split, or an explanation for one prediction.
arXiv:2607. 20543v1 Announce Type: cross Abstract: Reinforcement learning with verifiable rewards (RLVR) can improve one-sample accuracy while making a model worse under repeated sampling.
The paper investigates how approximate numerical solvers used in recursive state estimation can be repaired using bounded corrections, characterizing when such corrections meet local admissibility tolerances and how they influence finite‑horizon covariance. It derives error identities that separate solve error from gain drift, revealing quartic and sixth‑order contributions to the covariance response. The framework is applied to a power‑grid tolerance study, showing that learned corrections reduce the required conjugate‑gradient iterations, and it demonstrates a unified interface for classical, quantum, and hybrid solvers.
Deep equilibrium models promise input-adaptive implicit computation: harder problems should demand more solver iterations, and the solved equilibrium should encode the result of genuine iterative inference. We report a cautionary study of a port-Hamiltonian DEQ with a learned initialization on two reasoning tasks -- ProofWriter entailment over frozen DeBERTa embeddings and a BFS-verified graph-reachability benchmark -- in which the implicit computation is a silent no-op.