SiLR introduces a structure‑preserving admission and process reward mechanism for large language model (LLM) tool agents. Unlike traditional scalar‑score gates that can trap agents in plateau trajectories, SiLR shadow‑executes each proposal and admits it based on a product order over branch‑level violation states, ensuring safe and recoverable actions. Experiments on Gym‑ANM and CityLearn benchmarks show SiLR consistently recovers all multi‑action episodes and outperforms scalar gates, while also providing a robust reward signal for policy learning.
By Chenyu Zhou, Qiliang Jiang, Shuning Wu, Xu Zhou
arXiv:2607. 17240v1 Announce Type: new Abstract: When does a committed intermediate stage in an LLM reasoning pipeline earn its cost?
By Honglin Li (ShanghaiTech University)
arXiv:2606. 31023v1 Announce Type: cross Abstract: Hard-constrained sequential decision systems have no certified way to spend the test-time compute of modern AI: executing the multi-step drafts of a learned policy or a frozen LLM forfeits the feasibility guarantee a trusted solver provides, while invoking the solver at every step forfeits the speed the AI offers.
By Chenyu Zhou, Qiliang Jiang, Shuning Wu, Xu Zhou
arXiv:2607. 10972v1 Announce Type: new Abstract: Many evaluations of model outputs rely either on contracts checkable at evaluation time or on feedback that arrives within the operating loop.
By Aleh Manchuliantsau
Many evaluations of model outputs rely either on contracts checkable at evaluation time or on feedback that arrives within the operating loop. We study the complementary setting in which ground truth is delayed, censored, or private, so deterministic code cannot check correctness at scoring time and must instead issue a code-owned provisional forecast.
arXiv:2608. 14761v1 Announce Type: cross Abstract: At a finite public-chance cut, counterfactual regret minimization (CFR) must choose how many outcomes to evaluate before each regret update.
By Jiaxing Guo, Lei Ye
arXiv:2607. 29400v1 Announce Type: new Abstract: A routing decision can be revised at the next transaction, but a latched source exclusion persists across later decisions.
By Xiyang Zhang, Hongzhi Wang, Yuanhe Tian
arXiv:2609.22243v1 Announce Type: new
Abstract: Input-conditioned neural interventions raise a runtime question: what persists when one behavioral specification admits multiple actions whose validity...
By Xianliang Zeng, Zhanzhan Zhao
The study demonstrates that governance records—structured logs linking task contracts, model attempts, verifier decisions, and outputs—can serve as effective supervision for bounded AI models. Using a verifier-selected self‑training approach, the authors show that a Qwen3‑14B model trained on plans accepted by an independent VAL verifier achieved significant gains in plan acceptance across numerous PlanBench replanning cases, outperforming other selection strategies. The results highlight the feasibility of one‑shot execution and cumulative learning without relying on oracle targets or stronger teachers.
By Jesus Salas
arXiv:2607. 09709v1 Announce Type: new Abstract: Post-training a code generator against a learned judge can optimize proxy features that raise the score without improving the artifact.
By Chenyu Zhou, Qiliang Jiang, Shuning Wu, Xu Zhou
arXiv:2606. 25449v1 Announce Type: cross Abstract: A language model's memory can be worse than having no memory at all.
By Alex Kwon
arXiv:2608. 07583v1 Announce Type: cross Abstract: Multi-agent LLM systems route among model-backed advisors, yet a deployer rarely knows before shipping whether routing will help at all.
By Anchen Sun, Kaiqi Yang