arXiv:2609.09001v1 Announce Type: new
Abstract: Multi-step LLM reasoning lacks a machine-recheckable ledger: discarded reasoning paths leave no auditable record. We propose the Deposon scattering lay...
By Qihao Yuan
arXiv:2608. 08282v1 Announce Type: new Abstract: Tool-using language-model agents face constraints whose meaning changes with observations and prior actions.
By Ibne Farabi Shihab, Md Najmus Swaqeeb, Abu Sa-Adat Mohamed Moon-Im Al Ahsan
arXiv:2608. 03219v1 Announce Type: new Abstract: Benchmark gains are often treated as evidence of greater LLM capability.
By Yanchao Li, Wanhao Liu, Jiaqing Xie, Ben Gao, Yanbo Wang, Tianfan Fu, Yuqiang Li
Benchmark gains are often treated as evidence of greater LLM capability. Yet the same gain can reflect different changes in model behavior.
The paper introduces EvoResearcher, a training‑free, inference‑time protocol that enables a frozen large language model to perform cost‑bounded self‑reflection and early stopping. By iterating through generate → self‑critique → revise steps until a maximum depth or a CONFIRMED sentinel is reached, the model can self‑verify its answers within a strict compute budget. The protocol incorporates four self‑reflective meta‑reward components—correctness, efficiency, reflection depth, and tool‑call diversity—implemented as prompt‑level mechanisms, and is validated on Big‑Bench Hard, GSM8K, and MATH benchmarks, achieving comparable accuracy while terminating 82‑88% of items early with only about 2.1 generations per question.
By Wei Yu, Suxing Liu, Minjie Yu, Jiahao Wang, Zhijian Zheng, Haocheng Deng, Bing Li
arXiv:2606. 25449v1 Announce Type: cross Abstract: A language model's memory can be worse than having no memory at all.
By Alex Kwon
A language model's memory can be worse than having no memory at all. Give a model a memory that kept a wrong conclusion but dropped the work behind it, and it emits that stale value as a confident answer; give the same model an empty memory and it abstains.
arXiv:2607. 14137v2 Announce Type: cross Abstract: To answer a question about a program, move the program to where the question is decidable.
By Christoph Kirsch
arXiv:2608.20659v1 Announce Type: new
Abstract: Methods that transfer predictions from two-dimensional foundation models into three-dimensional segmentation are commonly grouped by task or representa...
By Wentao Sun, Yiping Chen, John S. Zelek, Jonathan Li
arXiv:2607. 05397v1 Announce Type: cross Abstract: Agent systems increasingly execute rather than advise.
By James Rhodes, George Kang
Grounded Continuation introduces a runtime verifier that classifies each utterance in an LLM conversation into one of eight epistemic operations and uses a symbolic engine to maintain a dependency map of claims and their supports. The verifier checks whether a new continuation is grounded by walking this map, a linear-time process that requires no additional LLM calls. On benchmarks such as ReviseQA and MemoryAgentBench, the verifier improves single-hop accuracy for several QA models, even enabling a 7B model to outperform GPT‑4o when guided by the verifier.
By Qisong He, Jinwei Hu, Xinmiao Huang, Changshun Wu, Yi Dong, Xiaowei Huang
The paper introduces a new evaluation protocol called checkpoint handoff to disentangle the contributions of reaching a target state and solving the task in reinforcement learning agents. By cloning states reached by one checkpoint and handing them to another without retraining, the authors separate the REACH metric (how often a policy arrives at a state confirmed to be a fixed number of actions from success) from the SOLVE metric (how often it finishes from that identical state). Across two benchmarks and pipelines, the analysis shows that RL history benefits RL solvers more than SFT solvers, and that independent REACH and SOLVE gaps predict overall performance.
By Xuan Liu, Jingbin Qian