Evo-Harness: Context-to-Harness Skill Compilation for Self-Evolving Agents
arXiv:2608. 15071v1 Announce Type: new Abstract: Learning from experience is critical for developing capable, self-improving large language model (LLM) agents.
Combee is a new framework that scales prompt learning for self‑improving language model agents by enabling many agents to run in parallel while learning from their combined traces. It uses parallel scans, an augmented shuffle mechanism, and a dynamic batch size controller to maintain quality and reduce delay. Experiments on AppWorld, Terminal‑Bench, Formula, and FiNER show up to 17× speedup over prior methods with comparable or better accuracy at similar cost.
arXiv:2608. 15071v1 Announce Type: new Abstract: Learning from experience is critical for developing capable, self-improving large language model (LLM) agents.
arXiv:2602. 11351v2 Announce Type: replace Abstract: Proactive large language model (LLM) agents aim to actively plan, query, and interact over multiple turns, enabling efficient task completion beyond passive instruction following and making them essential for real-world, user-centric applications.
arXiv:2606. 31270v1 Announce Type: cross Abstract: Computer-use agents, which leverage multimodal large language models (MLLMs) to operate computers and complete tasks, have attracted significant attention for their utility and versatility.
arXiv:2607. 20668v1 Announce Type: cross Abstract: TextGrad improves language-model systems by revising text from feedback.
Naive Prompt Optimization (NPO) is a lightweight, single‑lineage method that iteratively refines prompts using a teacher model’s rollout feedback. It matches or surpasses the performance of more complex optimizers like GEPA while requiring fewer rollouts, and its advantage grows with stronger teacher models. In interactive games, NPO remains competitive, and prompts optimized by NPO transfer well to other student models within the same family.
arXiv:2606. 05684v1 Announce Type: new Abstract: A central challenge for language agents is utilizing past experience to adapt to dynamic test-time conditions.
arXiv:2606. 03841v1 Announce Type: new Abstract: Recent progress in Large Language Model (LLM) agents has enabled promising advances in automated data science.
arXiv:2606. 00135v1 Announce Type: cross Abstract: Tool-calling is a central component of modern large language model (LLM) agents, equipping them with skills beyond their parametric knowledge.
arXiv:2511. 20297v2 Announce Type: replace Abstract: Large Language Model (LLM)-based agents are increasingly capable of complex, multi-step tasks such as GUI automation, tool use, and data manipulation, yet they cannot learn from experience: each new session rediscovers solutions from scratch.
The paper introduces Strategy Accumulation and Guided Execution (SAGE), a two-stage framework that makes automated fine-tuning of large language models cumulative. In the first stage, a multi-agent pipeline uses Monte Carlo Tree Search to explore training strategies while a Distillation Agent records task-specific insights and cross-task confidence scores into a structured repository. In the second stage, SAGE retrieves relevant experience from this repository to guide training on new tasks, achieving a 12.4‑percentage‑point improvement over a baseline pipeline without accumulated experience on nine unseen tasks.
The paper introduces Context Language Models (CLMs), which treat context as a mutable file that the model can update freely, enabling the model to learn what information to retain. CLMs built zero‑shot from existing models outperform state‑of‑the‑art context‑management methods on several benchmarks, achieving higher accuracy with fewer FLOPs. The authors also demonstrate that CLMs can be steered via natural‑language instructions and online reinforcement learning, and they propose a suffix‑cache reuse strategy that further reduces server‑side compute.
DeepEdu‑v1 is an AI‑tutoring system tailored for Vietnamese education that addresses data‑sovereignty and local curriculum alignment issues. It uses a long‑context inference engine to reduce retrieval calls and prefill latency by about 35%, and a self‑improving agentic layer that curates verified local knowledge without fine‑tuning. In deployment, DeepEdu achieves nearly twice the speed of standard vLLM serving and raises agentic accuracy from 70.0% to 79.5% on complex tasks, especially in financial reasoning and interactive‑agent benchmarks.