arXiv AI By Haochen Luo, Yi Huang, Sichun Luo, Fengyuan Liu, Lei Li, Zefa Hu, Junlan Feng, Qi Liu

Harness-Aware Self-Evolving: Co-Evolving Model Weights, Harness, and Task Solutions

Read the original on arXiv AI →

arXiv:2607. 03935v1 Announce Type: new Abstract: Self-evolving frameworks usually optimize task solutions while treating the surrounding harness as fixed.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
3d ago

Learning from Research: Toward Lifelong Agent Harness Evolution

The paper introduces ScholarEvolve, a framework that evolves the software harness of language agents by automatically incorporating insights from recent research papers. It organizes harness improvements into functional modules, uses topic modeling to identify distinct strategies, and evaluates combinations to boost task performance. Experiments show significant gains on AppWorld and Tau2-Bench, raising Qwen3.5-27B completion rates from 49.6% to 63.6% and GPT-5.4-mini pass@1 from 72.7% to 81.9%.

By Jingbo Yang, Kwei-Herng Lai, Xiaowen Wang, Yaar Harari, Evgeniy Gabrilovich, Shiyu Chang
arXiv AI
Jul 24

Self-Evolving Recommendation System: End-To-End Autonomous Model Optimization With LLM Agents

arXiv:2602. 10226v2 Announce Type: replace-cross Abstract: Optimizing large-scale machine learning systems, such as recommendation models for global video platforms, requires navigating a massive hyperparameter search space and, more critically, designing sophisticated optimizers, architectures, and reward functions to capture nuanced user behaviors.

By Haochen Wang, Yi Wu, Daryl Chang, Li Wei, Lukasz Heldt
arXiv AI
3d ago

Self-Evolving Harness on Multiple Tasks with the Agent as Its Own Optimizer

The paper introduces a self‑evolving harness framework where a frozen language‑model agent first solves tasks and then edits its own harness based on run records. Using a 49‑line seed harness, the evolved harness improves average scores on in‑distribution benchmarks by 4.48 points and on out‑of‑distribution benchmarks by 12.64 points, surpassing Codex on the former and matching it on the latter. Continued evolution on a specific out‑of‑distribution benchmark further raises performance, and the study analyzes emergent mechanisms such as output truncation and history compaction.

By Qiankai Xu