SelfSearch: Reward-Free Search for Self-Improving Agents
arXiv:2609. 37968v1 Announce Type: new Abstract: Advances in the coding capabilities of LLM agents allow them to inspect and modify their own instructions, tools, and execution procedures.
arXiv:2607. 14004v1 Announce Type: new Abstract: Most reported gains from agent-optimization methods are one-shot: an agent is optimized against a fixed benchmark and the resulting improvement is reported as if it were a stable property of the method.
arXiv:2609. 37968v1 Announce Type: new Abstract: Advances in the coding capabilities of LLM agents allow them to inspect and modify their own instructions, tools, and execution procedures.
arXiv:2608. 08466v1 Announce Type: new Abstract: Modern LLM agents are often improved by modifying prompts, tools, or workflows manually, while the executable scaffold surrounding the model---the \emph{harness}---is typically treated as a fixed artifact after deployment.
ActiveSaddler introduces automated curriculum learning for harness optimization, treating the evolving training curriculum as a non‑stationary bandit problem. It identifies reusable failure patterns, estimates learning progress for each, and balances revisiting known weaknesses with exploring new scenarios, allowing the curriculum to co‑evolve with the harness. Experiments on GAIA2 and Terminal‑Bench 2.0 show consistent improvements in harness performance, with Pass@1 gains of 4.4 and 7.5 percentage points over fixed‑order baselines.
arXiv:2607. 05297v1 Announce Type: new Abstract: Recent LLM agents tackle increasingly long-horizon, open-ended tasks, and external skills, reusable procedural knowledge supplied to the agent, further extend this capability.
EVOHARNESSBENCH is a new benchmark that tests how LLM-based agents handle changes in their tool, skill, and agent harnesses over time. It includes 17 deterministic harness streams with 802 tasks, 520 tools, 42 skills, and 62 agents, and evaluates agents in two settings: deployment evaluation and self‑evolving adaptation evaluation. The study finds that harness expansion can cause forgetting, adaptation gains are inconsistent, and preserving old competence does not always aid new capability adaptation, highlighting harness evolution as a distinct challenge for agent development.
The paper introduces a self‑evolving harness framework where a frozen language‑model agent first solves tasks and then edits its own harness based on run records. Using a 49‑line seed harness, the evolved harness improves average scores on in‑distribution benchmarks by 4.48 points and on out‑of‑distribution benchmarks by 12.64 points, surpassing Codex on the former and matching it on the latter. Continued evolution on a specific out‑of‑distribution benchmark further raises performance, and the study analyzes emergent mechanisms such as output truncation and history compaction.
arXiv:2609.04280v2 Announce Type: replace-cross Abstract: Modern LLM-based agents operate through a harness of tools, reusable skills, and specialist agents that shapes what they observe and what the...
arXiv:2608. 09629v1 Announce Type: new Abstract: Self-evolving agents are usually built around prescribed optimization pipelines: the framework decides how to gather evidence, revise a persistent artifact, select candidates, and stop.
arXiv:2609.24974v1 Announce Type: cross Abstract: Agent harnesses, the external systems that mediate model-environment interaction, can substantially improve agent performance, but their gains remain...
arXiv:2607. 14408v1 Announce Type: new Abstract: A self-evolving agentic loop repeatedly proposes a tweaked version of an agent (its prompt template or program) and accepts or rejects the change based on a per-iteration quality signal.
The paper introduces Regularized Recursive Self-Improvement of Agent Harnesses (RRSI), a method that applies regularization principles to the iterative editing of an LLM agent’s harness—prompts, control flow, tooling, memory, and context management. RRSI limits the number of edits per candidate, encourages novel trajectories, and uses a critic and pruner to filter out benchmark‑specific or ineffective changes, thereby favoring reusable agent mechanisms. Experiments on eight benchmarks show RRSI improves performance by up to 14.1 points on the training split and 4.7 points on out‑of‑distribution tests, while reducing policy token usage by 30% compared to unregularized evolution.
arXiv:2606. 25207v1 Announce Type: new Abstract: Hyperparameter Optimization (HPO) is essential for maximizing machine learning model performance, and its core challenge is sample efficiency: finding strong configurations within a limited budget.