arXiv:2607. 12790v1 Announce Type: new Abstract: Self-evolving agent systems improve by creating, revising, and retiring their own skills, but every such loop rests on a hidden assumption: a reliable evaluation metric already exists.
By Xing Zhang, Guanghui Wang, Yanwei Cui, Ziyuan Li, Wei Qiu, Bing Zhu, Peiyang He
The paper critiques the common practice of evaluating large‑language‑model (LLM) evolutionary search methods using a single seed and fixed iteration budget, arguing that this approach is insufficient. By testing three search strategies across five optimization tasks and varying both the number of seeds (width) and iterations (depth), the authors find that optimal budget allocation depends on the strategy, task, and total budget, and that strategy rankings shift with different budgets. They propose a measurement protocol that maps the seeds‑by‑iterations frontier and offers practical guidance for researchers.
By Tal Oved, Roi Pony, Oshri Naparstek, Udi Barzelay
arXiv:2608. 01378v1 Announce Type: new Abstract: Design campaigns in chemistry, materials science, and machine learning share a bottleneck: determining how good a candidate truly is requires an expensive evaluation - an experiment, a first-principles simulation, or a full training run.
By Shuangxiu (Max), Ma (Zachary), Wenhe (Zachary), Zhao
Self-evolving agent systems improve by creating, revising, and retiring their own skills, but every such loop rests on a hidden assumption: a reliable evaluation metric already exists. In many real applications it does not.
Agentic systems have widened the gap between producing candidate outputs and reviewing them. This paper asks a practical architectural question: should domain specialization be built into an evaluator's weights, or into the rule that decides when its judgment can be trusted?
The paper introduces Enrich‑Retrieve‑Rank, a scalable method for discovering capabilities in large agent ecosystems. It replaces in‑context routing with an offline enrichment step that converts sparse metadata into searchable profiles, followed by an online retrieve‑then‑rank pipeline that returns a ranked shortlist without invoking candidates. Experiments show that as the number of capabilities grows from 10 to 7,278, the new approach maintains higher top‑1 accuracy and reduces cost by 70× compared to full‑context baselines.
By Nazib Sorathiya, Daniel Zhang, Bardiya Akhbari