Optimize Cheap, Deploy Strong: Cost-Aware Cross-Tier Transfer for Evolutionary Optimization
arXiv:2608. 10694v1 Announce Type: cross Abstract: Evolutionary optimization of LLM prompts and agentic programs (e.
arXiv:2608. 16207v1 Announce Type: new Abstract: Consider a firm that surveys its competition for a particular agentic task and seeks to offer superior accuracy at every competitor price point.
arXiv:2608. 10694v1 Announce Type: cross Abstract: Evolutionary optimization of LLM prompts and agentic programs (e.
The paper introduces ARTEMIS, a no-code evolutionary optimization platform that automatically tunes large language model (LLM) agents by jointly optimizing prompts, tool descriptions, and parameters using semantically-aware genetic operators. Starting from a benchmark script and natural language goals, ARTEMIS discovers configurable components, extracts performance signals from execution logs, and evolves configurations without architectural changes. Experiments on four agent systems show significant gains: a 13.6% increase in acceptance rate for the ALE Agent, a 10.1% performance boost for the Mini‑SWE Agent, a 36.9% token‑reduction for the CrewAI Agent, and a 22% accuracy improvement for the MathTales‑Teacher Agent using a smaller open‑source model.
The paper introduces Regularized Recursive Self-Improvement of Agent Harnesses (RRSI), a method that applies regularization principles to the iterative editing of an LLM agent’s harness—prompts, control flow, tooling, memory, and context management. RRSI limits the number of edits per candidate, encourages novel trajectories, and uses a critic and pruner to filter out benchmark‑specific or ineffective changes, thereby favoring reusable agent mechanisms. Experiments on eight benchmarks show RRSI improves performance by up to 14.1 points on the training split and 4.7 points on out‑of‑distribution tests, while reducing policy token usage by 30% compared to unregularized evolution.
EarlyEval introduces a lightweight framework that predicts an LLM agent’s final outcome early in its execution, allowing the run to halt when a LightGBM classifier reaches a calibrated confidence threshold. By training success and failure classifiers on behavioral, textual, and reference-solution features, EarlyEval can cut 13%-26% of agent steps and up to 44.1% of input tokens while maintaining 89%-97% prediction accuracy. Across three benchmarks—SWE-bench Verified, TerminalBench, and Toolathlon—this approach reduces evaluation costs with minimal impact on per-agent resolve rates.
arXiv:2607. 14408v1 Announce Type: new Abstract: A self-evolving agentic loop repeatedly proposes a tweaked version of an agent (its prompt template or program) and accepts or rejects the change based on a per-iteration quality signal.
arXiv:2607. 20468v1 Announce Type: new Abstract: AI agents are increasingly used to automate research and development tasks, yet existing benchmarks typically evaluate them on prescribed workflows or narrow action spaces.
arXiv:2607. 09600v1 Announce Type: new Abstract: Enhancing the reasoning capabilities of large language model (LLM) agents requires effective orchestration of diverse expert models and tools.
arXiv:2605. 28556v2 Announce Type: replace Abstract: As agent capabilities advance, existing benchmarks, such as $\tau^2$-Bench, are becoming increasingly saturated.
arXiv:2606. 26294v1 Announce Type: cross Abstract: Self-improving agents are state-of-the-art (SOTA) on agentic coding benchmarks and have recently been extended to general domains.
arXiv:2609.38445v1 Announce Type: new Abstract: Frontier LLMs are increasingly used to automate scientific research through iterative search. We distinguish idea-driven search from solution-driven se...
arXiv:2608. 05519v1 Announce Type: new Abstract: Agent benchmarks usually measure task completion and treat resource use as an auxiliary statistic.
arXiv:2607. 29241v1 Announce Type: cross Abstract: Optimizing modern recommender models still depends heavily on engineers manually iterating over architectural, objective, and training-strategy changes.