arXiv:2606. 18617v1 Announce Type: cross Abstract: There exist numerous tutor training platforms.
By Danielle R. Thomas, Marie Cynthia Abijuru Kamikazi, Clara Brandt, Conrad Borchers, Kenneth R. Koedinger
The paper introduces StepLearn, a nonparametric framework for prequential test‑time learning in large language model agents. StepLearn separates immediate use of informative transitions from persistent trust, turning each transition into a hypothesis that guides the next step and only reusing it after prospective validation across episodes. Experiments on WebArena‑Lite and ALFWorld show StepLearn improves success rates by 2.2–12.7 percentage points over the strongest baseline, with benefits evident from the first task attempts.
By Tong Zhao, Reed Li, Yuyang Hu, Yutao Zhu, Haijin Liang, Haibo Shi, Yu Lu, Zhicheng Dou
arXiv:2609.05435v2 Announce Type: replace
Abstract: Can language agents continually learn from experience, turning earlier interactions into reusable capabilities? AhaBench evaluates this ability thr...
By Zerui Cheng, Jiawei Xu, Huacan Chai, Jiayang Sun, Pramod Viswanath, Maxm Pan
arXiv:2602. 13241v3 Announce Type: replace-cross Abstract: Emergency call-takers form the first operational link in public safety response, handling over 240 million calls annually while facing a sustained training crisis: staffing shortages exceed 25\% in many centers, and preparing a single new hire can require up to 720 hours of one-on-one instruction that removes experienced personnel from active duty.
By Zirong Chen, Meiyi Ma
arXiv:2605.08693v3 Announce Type: replace
Abstract: Skills provide an effective mechanism for improving LLM agents on complex tasks, yet in existing agent frameworks, their creation, refinement, and...
By Min Yang, Jinghua Piao, Xu Xia, Xiaochong Lan, Jiaju Chen, Yongshun Gong, Yong Li
arXiv:2607. 21419v1 Announce Type: new Abstract: In long-horizon LLM agent reinforcement learning, weak policies often repeat similar failures, producing uninformative rollout trajectories and limiting effective policy optimization.
By Yipeng Shi, Zhipeng Ma, Yue Wang, Qitai Tan, Yang Li, Peng Chen, Zhengzhou Zhu