arXiv:2606. 19808v1 Announce Type: new Abstract: Test-time reasoning is increasingly used as a serving-time control knob, but extra reasoning is not uniformly valuable: it can repair failed attempts, waste compute on already-correct answers, or introduce harmful answer changes.
By Sajib Acharjee Dip, Dawei Zhou, Liqing Zhang
The paper investigates whether self-consensus—stopping a reasoning model when its partial trajectory’s answers agree—can safely reduce inference cost. A large sweep of 3,520 consensus rules on two models and three benchmarks failed to meet predefined safety criteria, while a boundary‑confidence control (DEER) succeeded. The study shows that agreement signals that an answer persists under a fixed probing procedure, not that reasoning has finished, leading to premature stops and missed corrections even when token savings are significant.
By Yunxiang Mo, Donghao Zhao, Hejia Geng
arXiv:2606. 15621v1 Announce Type: new Abstract: Per-token counterfactual credit estimation asks which token in a language-model rollout caused the final answer to be right or wrong: cut the transcript at a pivot, substitute an alternative token, replay continuations, and compare outcomes.
By Nils Matteson
arXiv:2609. 10954v1 Announce Type: new Abstract: Continual world models must decide whether new data justify changing the model.
By Anqi Peter Li, Kaden Kim
arXiv:2608. 10145v1 Announce Type: new Abstract: LeWorldModel trains a latent world model with a prediction loss and a single anti-collapse regulariser, and reports approximately 87% of goals reached on TwoRoom, its simplest diagnostic environment.
By Joyjeet Singh
arXiv:2609.13714v1 Announce Type: new
Abstract: An updated model can improve an aggregate metric while degrading a slice that matters to a downstream user. We study checkpoint selection subject to no...
By Shengwei Zhang, Tao Wu, Fei Qian