The paper introduces StepLearn, a nonparametric framework for prequential test‑time learning in large language model agents. StepLearn separates immediate use of informative transitions from persistent trust, turning each transition into a hypothesis that guides the next step and only reusing it after prospective validation across episodes. Experiments on WebArena‑Lite and ALFWorld show StepLearn improves success rates by 2.2–12.7 percentage points over the strongest baseline, with benefits evident from the first task attempts.
By Tong Zhao, Reed Li, Yuyang Hu, Yutao Zhu, Haijin Liang, Haibo Shi, Yu Lu, Zhicheng Dou
arXiv:2605. 31191v2 Announce Type: replace Abstract: We investigate how teacher-student capacity relationships modulate knowledge distillation (KD) effectiveness in ResNet-based image classification on CIFAR-10.
By Umut Onur Yasar
arXiv:2606. 29280v1 Announce Type: cross Abstract: We identify intervention bias as a previously unquantified failure mode of zero-shot large-language-model (LLM) educational advisory agents: without task-specific training, they recommend action when a hindsight-optimal oracle policy mandates inaction.
By Craig Atkinson
arXiv:2606. 14239v1 Announce Type: new Abstract: Agent skills are structured procedural packages that guide frozen LLM agents in specialized workflows.
By Haowen Gao, Haoran Chen, Can Wang, Shasha Guo, Liang Pang, Zhaoyang Liu, Huawei Shen, Xueqi Cheng
arXiv:2609.05435v1 Announce Type: new
Abstract: Modern language agents are expected to operate over long horizons: they ask follow-up questions, reuse worked examples, handle tool feedback, and adapt...
By Zerui Cheng, Jiawei Xu, Huacan Chai, Jiayang Sun, Pramod Viswanath, Maxm Pan
arXiv:2609.37472v1 Announce Type: cross
Abstract: Behavioral tests measure how a language model reads evidence. We ask whether those measurements help choose a recommendation interface. We evaluate s...
By Han Chen, Yingrui Li
arXiv:2606. 15390v1 Announce Type: cross Abstract: LLM agents can improve without weight updates by accumulating natural-language skills from experience, but current systems entrust every decision about which skills to keep and how to apply them to LLM judgment alone.
By Yixuan Wang, Yiyang Zhou, Yiming Liang, Congyu Zhang, Fuxiao Liu, Jiawei Zhou, Huaxiu Yao
arXiv:2609.05435v2 Announce Type: replace
Abstract: Can language agents continually learn from experience, turning earlier interactions into reusable capabilities? AhaBench evaluates this ability thr...
By Zerui Cheng, Jiawei Xu, Huacan Chai, Jiayang Sun, Pramod Viswanath, Maxm Pan
The study examines how two small instruction‑tuned language models, Qwen2.5‑1.5B and Llama‑3.2‑1B, respond to user pushback on TriviaQA. When initially correct, the models flip to a wrong answer in about 42–43% of cases, with the effectiveness of different pushback styles varying by model. Attempts to decode capitulation from the pre‑response residual stream fail under a rigorous validation protocol, revealing overfitting and a measurement hazard that underestimates capitulation by 18–24 percentage points.
By Saad Aamir, Muhammad Awais Bin Adil
The study clusters 5,000 EdNet-KT3 learners into eight study‑strategy groups based on early‑session behaviors such as resource use, revision, video watching, and problem practice. These clusters predict later engagement metrics—like continued practice and session completion—but do not reliably forecast later unassisted accuracy or mastery. The findings suggest that behavioral clustering captures learning styles and engagement patterns rather than knowledge gains.
By Qingchuan Lyu, Yingxin Li, Albert Yang
arXiv:2608. 09826v1 Announce Type: cross Abstract: Reinforcement learning with verifiable rewards yields no group-relative signal when rollout groups are uniformly correct or uniformly wrong, which account for 63.
By Yubo Jiang, Fengying Xie, Zhiguo Jiang, Haopeng Zhang
arXiv:2607. 07050v3 Announce Type: replace-cross Abstract: Top-K teacher logits make on-policy distillation tractable, but probability mass is not the same as decision support.
By Jiabin Shen, Guang Chen, Chengjun Mao