ACLArena: Agent Continue Learning in Multi-stage Post-training
Read the original on arXiv AI →The Flow has not summarised this story yet — read it at arXiv AI.
The Flow has not summarised this story yet — read it at arXiv AI.
The paper introduces Harness Continual Learning (HCL), a paradigm where an agent’s state evolves through prompts, memories, tools, skills, and routing rules while keeping the underlying foundation model frozen. HCL defines harness-level forgetting and proposes a guarded evolution process involving a Continual Optimizer and Evaluator to ensure improvements without losing prior behavior. Experiments across textual reasoning, multimodal perception, and open‑world interaction show over 10% performance gains and demonstrate how the stability–plasticity trade‑off can be explicitly tuned.
The paper introduces Harness Continual Learning (HCL), a paradigm where an agent’s state evolves through prompts, memories, tools, skills, and routing rules while keeping the foundation model frozen. HCL defines harness-level forgetting and proposes guarded harness evolution with a Continual Optimizer and Evaluator to balance improvement, retention, and validity. Experiments across textual reasoning, multimodal perception, and open‑world interaction show over 10% performance gains and demonstrate explicit control over the stability–plasticity trade‑off.
arXiv:2608. 03874v1 Announce Type: new Abstract: Modern agent frameworks equip large language models with external skill libraries to solve complex tasks.
arXiv:2609.05435v2 Announce Type: replace Abstract: Can language agents continually learn from experience, turning earlier interactions into reusable capabilities? AhaBench evaluates this ability thr...
Modern agent frameworks equip large language models with external skill libraries to solve complex tasks. However, it remains unclear whether these systems can effectively evolve their skills and whether the resulting skills improve task-solving capabilities.
arXiv:2607. 07847v1 Announce Type: new Abstract: As large language models (LLMs) become increasingly capable, the next question is how can we enable models to continually learn?