Evo-Harness: Context-to-Harness Skill Compilation for Self-Evolving Agents
arXiv:2608. 15071v1 Announce Type: new Abstract: Learning from experience is critical for developing capable, self-improving large language model (LLM) agents.
The paper introduces ScholarEvolve, a framework that evolves the software harness of language agents by automatically incorporating insights from recent research papers. It organizes harness improvements into functional modules, uses topic modeling to identify distinct strategies, and evaluates combinations to boost task performance. Experiments show significant gains on AppWorld and Tau2-Bench, raising Qwen3.5-27B completion rates from 49.6% to 63.6% and GPT-5.4-mini pass@1 from 72.7% to 81.9%.
arXiv:2608. 15071v1 Announce Type: new Abstract: Learning from experience is critical for developing capable, self-improving large language model (LLM) agents.
arXiv:2606. 07603v1 Announce Type: cross Abstract: Large language models (LLMs) exhibit strong reasoning capabilities, yet most LLM-based agents are statically deployed and unable to improve through task interactions.
arXiv:2603. 20667v2 Announce Type: replace-cross Abstract: Existing prompt-optimization techniques rely on local signals, causing poor generalization across tasks.
arXiv:2607. 26722v1 Announce Type: cross Abstract: Harness plays a critical role in large language model agent performance, and building a high-performing harness requires substantial expert effort.
The paper introduces a self‑evolving harness framework where a frozen language‑model agent first solves tasks and then edits its own harness based on run records. Using a 49‑line seed harness, the evolved harness improves average scores on in‑distribution benchmarks by 4.48 points and on out‑of‑distribution benchmarks by 12.64 points, surpassing Codex on the former and matching it on the latter. Continued evolution on a specific out‑of‑distribution benchmark further raises performance, and the study analyzes emergent mechanisms such as output truncation and history compaction.
Large Language Models (LLMs) have driven rapid progress in autonomous agents, yet standard evaluations remain confined to static task solving. An emerging frontier is harness evolution---the agent's capacity to autonomously optimize its own operating harness.
The paper introduces PACEvolve, a framework that improves self‑evolving agents powered by Large Language Models by addressing their tendency to become trapped in local contexts and repeat flawed hypotheses. It does so through three techniques: Hierarchical Context Management to prune memory, Momentum‑Based Backtracking to escape local minima, and a self‑adaptive Collaborative Evolution policy to balance refinement and knowledge transfer. These methods enable the agents to maintain a global view of search momentum and achieve state‑of‑the‑art results on complex evolutionary benchmarks.
arXiv:2606. 07412v1 Announce Type: cross Abstract: LLM-driven software engineering agents have become a central testbed for real-world language-model capability, yet their training remains limited by the availability of high-quality SWE tasks.
HarnessEvolve is a self‑evolving framework that improves agent harnesses—prompts, skills, tools, and execution logic—by learning from reference trajectories. It separates execution, evaluation, optimization, and gating into independent modules, addressing credit assignment failure, shortcut learning, and catastrophic forgetting. The approach uses reference trajectories to extract error signals, applies quality and performance gates to candidate updates, and validates updates on held‑out data, consistently outperforming state‑of‑the‑art baselines across diverse benchmarks.
arXiv:2608. 01918v1 Announce Type: new Abstract: Harness design plays a critical role in agent performance by shaping how large language models (LLMs) perceive, reason over, and act within executable environments.
arXiv:2605.24539v2 Announce Type: replace Abstract: Harness evolution enables frozen language model agents to adapt to unfamiliar tasks by modifying the external programs that govern their behavior....
arXiv:2604. 17406v3 Announce Type: replace Abstract: The convergence of large language models and agents is catalyzing a new era of scientific discovery: Agentic Science.