arXiv:2606.21399v2 Announce Type: replace
Abstract: Runtime oversight often intervenes when an LLM agent's calibrated failure score crosses a threshold. Yet states with the same failure risk can diff...
By Chubin Zhang, Zhenglin Wan, Xingrui Yu, Jingxuan Wu, Qi Wen, Pengfei Zhou, Wangbo Zhao, Ivor Tsang
The paper introduces StepLearn, a nonparametric framework for prequential test‑time learning in large language model agents. StepLearn separates immediate use of informative transitions from persistent trust, turning each transition into a hypothesis that guides the next step and only reusing it after prospective validation across episodes. Experiments on WebArena‑Lite and ALFWorld show StepLearn improves success rates by 2.2–12.7 percentage points over the strongest baseline, with benefits evident from the first task attempts.
By Tong Zhao, Reed Li, Yuyang Hu, Yutao Zhu, Haijin Liang, Haibo Shi, Yu Lu, Zhicheng Dou
arXiv:2607. 06503v1 Announce Type: new Abstract: Large language model (LLM) agents solving multi-step tasks frequently commit to trajectories that are doomed to fail, yet continue to consume substantial inference compute before the failure becomes observable.
By Kai Ruan, Zihe Huang, Ziqi Zhou, Qianshan Wei, Xuan Wang, Hao Sun
Large language model (LLM) agents solving multi-step tasks frequently commit to trajectories that are doomed to fail, yet continue to consume substantial inference compute before the failure becomes observable. We show that failure is predictable early from the agent's internal representations: lightweight per-round probes on hidden activations anticipate eventual episode failure as early as the first interaction round, where scorers reading only the agent's observable behavior are barely better than chance.
arXiv:2607. 05458v1 Announce Type: cross Abstract: Large language model (LLM) agents are usually improved by changing prompts, models, or hand-written workflows, while the execution harness around the model is treated as fixed infrastructure.
By Haiwen Yi, Xinyuan Song
arXiv:2608. 03222v1 Announce Type: cross Abstract: Software engineering (SWE) agents resolve repository-level issues through long trajectories that grow increasingly expensive as context accumulates.
By Chenyu Wang, Yunbo Lyu, Junda He, Zhou Yang, Chenxing Zhong, Yaniv Harel, David Lo
arXiv:2606. 01311v1 Announce Type: cross Abstract: Large language model (LLM) agents increasingly rely on reusable external skills to solve long-horizon interactive tasks.
By Zhuoyun Yu, Xin Xie, Wuguannan Yao, Chenxi Wang, Lei Liang, Xiang Qi, Shumin Deng
The paper introduces STEPGATE, an uncertainty‑aware handoff framework that evaluates each step of a small language model (SLM) agent and selectively escalates difficult steps to a stronger model. On a 52‑task single‑step benchmark, the Qwen2.5‑1.5B/7B pair achieved 82.7% task success with only 30.8% escalation, outperforming local‑only and random escalation baselines. In multi‑turn tests, STEPGATE reached 69.0% trajectory success and 84.0% action success while using only 30.0% cloud actions, demonstrating that step‑level escalation can close much of the performance gap to a stronger backend with fewer remote tokens.
By Abolfazl Younesi
Large language model (LLM) agents are usually improved by changing prompts, models, or hand-written workflows, while the execution harness around the model is treated as fixed infrastructure. We argue that this harness is itself a learnable control layer.
arXiv:2609.14636v1 Announce Type: new
Abstract: On-policy distillation (OPD) has become a standard approach for transferring capabilities from large teachers to compact students. Its cost, however, i...
By Zhiyu Gui, Kexin Huang, Jia Guo, Junkang Wu, Zihao Wang, Zhiqiang Zhang, Jun Zhou, Jiancan Wu, Xiang Wang
Software engineering (SWE) agents resolve repository-level issues through long trajectories that grow increasingly expensive as context accumulates. Failed runs tend to be longer and exhibit redundant exploration or looping, suggesting that some failures may be detectable before completion.
The paper introduces checkpoint handoff, an evaluation protocol that separates an agent’s ability to reach useful states from its ability to complete tasks in reinforcement learning. By using one checkpoint as a reacher up to a handoff point and another as a solver from the same replayed history, the authors can measure Reach (how often states within a fixed number of actions from success are achieved) and Solve (how often the task is completed from those states). Experiments on TravelPlanner and ALFWorld show that switching the solver from supervised fine‑tuning to RL yields larger gains when RL is used as the reacher, indicating that RL more effectively finds solvable states.
By Xuan Liu, Jingbin Qian