The paper investigates Training‑Inference Mismatch (TIM) in large‑language‑model reinforcement learning, where rollout generation and policy optimization produce differing token probabilities despite identical model weights. By creating a zero‑mismatch diagnostic setting called VeXact, the authors isolate TIM and demonstrate that even minor token‑level numerical disagreements can trigger training collapse. They further show that TIM alters the effective optimization problem and propose remedies to mitigate its impact.
By Tianle Zhong, Neiwen Ling, Yifan Pi, Zijun Wei, Tianshu Yu, Geoffrey Fox, Peng Wu, Xiao Yu
REVERSAL-BENCH is a benchmark that introduces a continuous reversibility parameter ρ∈[0,1] and a reset oracle to evaluate how well reinforcement learning agents can recover from irreversible states across eight manipulation tasks in five physics engines. Experiments show a sharp reversibility cliff: reset‑free agents become trapped in irrecoverable states as ρ increases, while episodic agents continue learning steadily. The benchmark also provides a large multi‑simulator dataset and demonstrates that safety shields can predict recoverability but only succeed when the agent can avoid the trap.
By Riyaaz Shaik, Chandru Venkataraman
arXiv:2608. 05702v1 Announce Type: new Abstract: Scientific machine learning commonly validates models at the level of a subdomain, a benchmark split, or an explanation for one prediction.
By Gnankan Landry Regis N'guessan, Bum Jun Kim
arXiv:2607. 20543v1 Announce Type: cross Abstract: Reinforcement learning with verifiable rewards (RLVR) can improve one-sample accuracy while making a model worse under repeated sampling.
By Todd Zhou
The paper investigates how approximate numerical solvers used in recursive state estimation can be repaired using bounded corrections, characterizing when such corrections meet local admissibility tolerances and how they influence finite‑horizon covariance. It derives error identities that separate solve error from gain drift, revealing quartic and sixth‑order contributions to the covariance response. The framework is applied to a power‑grid tolerance study, showing that learned corrections reduce the required conjugate‑gradient iterations, and it demonstrates a unified interface for classical, quantum, and hybrid solvers.
By Yanjun Ji, Dennis Willsch, Orkun \c{S}ensebat, Priyanka Arkalgud Ganeshamurthy, Zhi Pei, M. Sahnawaz Alam, Ivelina Stoyanova, Frank K. Wilhelm, Bo Zhao, Chao Wang, Kristel Michielsen
Deep equilibrium models promise input-adaptive implicit computation: harder problems should demand more solver iterations, and the solved equilibrium should encode the result of genuine iterative inference. We report a cautionary study of a port-Hamiltonian DEQ with a learned initialization on two reasoning tasks -- ProofWriter entailment over frozen DeBERTa embeddings and a BFS-verified graph-reachability benchmark -- in which the implicit computation is a silent no-op.
The paper introduces TRACE, a digital‑advertising diagnostic environment that uses simulated interventions to generate verifiable rewards for training reasoning agents. By injecting controlled interventions into a simulator, the hidden cause of anomalies becomes an oracle label, enabling agents to learn to identify root causes and affected segments through noisy, confounded evidence. Experiments show that reinforcement learning with these synthesized rewards outperforms large prompted baselines, achieving higher accuracy while using fewer tool calls.
By Rui Sun, Zhan Shi, Bing He
The paper introduces ActObs, a supervised fine‑tuning method that, unlike standard approaches, also predicts environment observations in agent trajectories. While both ActObs and action‑only training perform similarly after initial fine‑tuning, ActObs diverges during subsequent reinforcement learning, yielding higher pass@k scores on several benchmarks and better cross‑domain task performance. The authors attribute this advantage to ActObs’s joint supervision, which preserves observation gradients and prevents the policy from over‑specializing on actions alone.
By Juzheng Zhang, Disha Makhija, Manoj Ghuhan Arivazhagan, Vinayshekhar Bannihatti Kumar, Rashmi Gangadharaiah
arXiv:2608. 04788v1 Announce Type: cross Abstract: Large language model agents are commonly trained through reinforcement learning with sparse trajectory-level rewards, which offer limited guidance on how strongly individual tokens should be updated.
By Yi Yang, Cong Qin, Xiaodan Liu, Chishui Chen, Qing Dong, Yan Zhang, Cao Liu, Zhao Yang, Lu Pan, Jiaye Lin, Yi Feng
arXiv:2607. 11116v1 Announce Type: cross Abstract: Deep equilibrium models promise input-adaptive implicit computation: harder problems should demand more solver iterations, and the solved equilibrium should encode the result of genuine iterative inference.
By Joyjeet Singh
arXiv:2601. 15141v2 Announce Type: replace Abstract: Agentic Reinforcement Learning (RL) has empowered Large Language Models (LLMs) to utilize tools like Python interpreters for complex problem-solving.
By Tianshi Xu, Yuteng Chen, Meng Li
arXiv:2607. 07379v1 Announce Type: new Abstract: In agentic scientific machine learning (SciML), large language model (LLM) agents can discover surrogate models and select one by an automated score, typically an error metric.
By Diab W. Abueidda, Bilal Ahmed, Panos Pantidis, Mostafa E. Mobasher