arXiv Machine Learning

AhaBench: Do Agents Turn Experience into Reusable Insights? A Long-Horizon Benchmark for Continual Learning

arXiv AI
1d ago

Learn Now, Use Next, Trust Later: Prequential Test-Time Learning for LLM Agents

The paper introduces StepLearn, a nonparametric framework for prequential test‑time learning in large language model agents. StepLearn separates immediate use of informative transitions from persistent trust, turning each transition into a hypothesis that guides the next step and only reusing it after prospective validation across episodes. Experiments on WebArena‑Lite and ALFWorld show StepLearn improves success rates by 2.2–12.7 percentage points over the strongest baseline, with benefits evident from the first task attempts.

By Tong Zhao, Reed Li, Yuyang Hu, Yutao Zhu, Haijin Liang, Haibo Shi, Yu Lu, Zhicheng Dou
arXiv AI
Aug 5

EduClaw-Bench: A Long-Horizon Benchmark for Pedagogical LLM Agents with Simulated Learners

arXiv:2608. 03206v1 Announce Type: cross Abstract: Large language models (LLMs) power educational applications from tutoring to essay scoring, but each is a point solution to a single task, and only recently have these point solutions been integrated into agents operating over a learning management system (LMS).

By Unggi Lee, Sookbun Lee, Yeil Jeong, Eunjoo Lee, Minchul Shin, Hoilym Kwon
arXiv AI
Sep 23

ACLArena: Agent Continue Learning in Multi-stage Post-training

arXiv:2609.23989v1 Announce Type: new Abstract: Building general-purpose agents for industrial deployment requires integrating multiple capabilities, each typically acquired at a distinct stage of tr...

By Haixin Wang, Xiaoxuan Wang, Junkai Zhang, Han Zhang, Renliang Sun, Alexander K Taylor, Yidan Shi, Haoran Deng, Chenguang Wang, Jason Cong, Yizhou Sun, Wei Wang
arXiv AI
Aug 20

Harness Continual Learning: Continual Adaptation Beyond Model Parameters

The paper introduces Harness Continual Learning (HCL), a paradigm where an agent’s state evolves through prompts, memories, tools, skills, and routing rules while keeping the underlying foundation model frozen. HCL defines harness-level forgetting and proposes a guarded evolution process involving a Continual Optimizer and Evaluator to ensure improvements without losing prior behavior. Experiments across textual reasoning, multimodal perception, and open‑world interaction show over 10% performance gains and demonstrate how the stability–plasticity trade‑off can be explicitly tuned.

By Borui Kang, Jinrui Gu, Junhan Lv, Wenbin Li, Lei Wang, Yang Gao
arXiv Computation and Language
Sep 7

EVOHARNESSBENCH: Can Your Agents Keep Pace with an Evolving Harness?

EVOHARNESSBENCH is a new benchmark that tests how LLM-based agents handle changes in their tool, skill, and agent harnesses over time. It includes 17 deterministic harness streams with 802 tasks, 520 tools, 42 skills, and 62 agents, and evaluates agents in two settings: deployment evaluation and self‑evolving adaptation evaluation. The study finds that harness expansion can cause forgetting, adaptation gains are inconsistent, and preserving old competence does not always aid new capability adaptation, highlighting harness evolution as a distinct challenge for agent development.

By Zixuan Ke, Vaidehi Patil, Haizhou Shi, Yang Li, Ye Liu, Sarath Shekkizhar, Anurag Koul, Jiayu Wang, Xuan Phi Nguyen, Semih Yavuz, Mohit Bansal, Shafiq Joty
arXiv AI
Jun 30

UCOB: Learning to Utilize and Evolve Agentic Skills via Credit-Aware On-Policy Bidirectional Self-Distillation

arXiv:2606. 29502v1 Announce Type: new Abstract: Skill memories can improve agentic reinforcement learning by reusing past experience as textual guidance, but retrieved skills are not oracular: they may help in one state while misleading the same policy in another.

By Songjun Tu, Chengdong Xu, Qichao Zhang, Yiwen Ma, Yaocheng Zhang, Linjing Li, Dong Li, Xiangyuan Lan, Dongbin Zhao