The paper introduces the concept of proof‑carrying cognition, aiming to close the verification gap in language‑model reasoning by using reality‑settled rewards. It presents a theoretical framework linking verifier‑gold correlation to compute‑capability trade‑offs, demonstrates that unsound verifiers degrade under best‑of‑N selection while sound verifiers improve, and proposes a new benchmark metric, Soundness‑under‑Pressure, for evaluating reality‑settled reasoning systems.
By Eshwar Reddy M, Sourav Karmakar
arXiv:2608. 15565v1 Announce Type: new Abstract: Experience-learning agents for optimization modeling improve by storing verified skills, but existing learners admit knowledge by checking against known answers, which real ticket streams do not provide.
By Junbo Jacob Lian, Huiling Chen, Hanzhang Qin, Chung-Piaw Teo
The paper introduces belief‑shift branching, a method for placing forks in tree‑structured reinforcement learning rollouts by identifying points where a model’s answer belief changes most. Unlike traditional structural or entropy‑based approaches, belief‑shift uses a probe, logit‑lens depth profile, or learned activation direction to locate pivots in the value curve, incurring minimal computational overhead. Experiments across multiple models and benchmarks show that belief‑shift forking consistently outperforms baseline methods, yielding significant gains in mathematics and code tasks.
The paper introduces Reward‑Informed Sparse Autoencoders (RI‑SAEs), which use reinforcement‑learning rewards to curate data for training sparse autoencoders on language‑model activations. On Llama‑3.1‑8B, a sparse subset of features separates high‑reward from low‑reward reasoning continuations, but control experiments show this separation largely reflects solution completeness rather than true reasoning quality. The authors conclude that reward filtering can cheaply reuse RL signals for interpretability, though most of the discovered features capture completion form rather than deep reasoning.
By Tanvi Nagilla, Alexander Jameson, Daniel Manta, Shayaan Uddin
arXiv:2607. 03436v1 Announce Type: new Abstract: Routing among large language models (LLMs) promises better quality at lower cost, motivated by the reported gap between learned routers and a per-instance oracle.
By Teng-Ruei Chen
arXiv:2608. 11560v1 Announce Type: new Abstract: Personalizing marketing messages with contextual multi-armed bandits (CMABs) drives real business value, yet the objective that ultimately matters - a downstream conversion - is observed only weeks later, too late to drive online learning.
By Sang Su Lee, Vineeth Loganathan, Shishir Dash, Vijay Raghavan
arXiv:2606. 10064v1 Announce Type: cross Abstract: Small-model agentic post-training is bottlenecked less by the algorithm than by the trajectory substrate it consumes.
By Shardul Bansal, Seth Schilbe, Jarrod Barnes
arXiv:2606. 25556v1 Announce Type: cross Abstract: Stepwise group-based RL is an attractive way to train long-horizon LLM agents without a learned critic: it reuses multiple sampled rollouts to estimate local advantages.
By Hanyang Wang, Weijieying Ren, Yuxiang Zhang, Ding Cao, Zhizhao Zeng, Ke Zeng, Tianxiang Zhao
arXiv:2603. 13356v2 Announce Type: replace Abstract: Robust reinforcement learning typically assumes that feedback sources are either globally trustworthy or corrupted within a fixed global budget.
By Majid Ghasemi, Mark Crowley
arXiv:2609.01274v1 Announce Type: new
Abstract: Reinforcement learning with verifiable rewards (RLVR) improves language-model reasoning, but how these gains relate to inference-time decoding and sear...
By Wenhe Sun, Cunxiang Wang, Zijun Yao, Yixin Cao
arXiv:2606. 21253v2 Announce Type: replace Abstract: Continual learning that is gradient-free, local, online, and append-only is attractive for edge and streaming deployment, but its value is usually argued informally.
By Jianwei Lou (RailMind Systems, Neuss, Germany)
The paper introduces the concept of Wide Learning, which examines how a learner’s internal state can expand its ability to generate informative evidence under fixed resources and primitive affordances. By formalizing effective epistemic reach—defined by learner state, deployment budget, reliability threshold, and evaluation distribution—the authors demonstrate, through a controlled construction, that learning can significantly alter the probability of successfully realizing a diagnostic that was previously unlikely. The study shows that even with identical observable laws, a calibrated learner can achieve perfect diagnostic realization, highlighting the impact of learning on the scope of attainable evidence.
By Junzhou Chen