arXiv AI

Evolution or Illusion? Rethinking Evaluation in LLM Evolutionary Search

The paper critiques the common practice of evaluating large‑language‑model (LLM) evolutionary search methods using a single seed and fixed iteration budget, arguing that this approach is insufficient. By testing three search strategies across five optimization tasks and varying both the number of seeds (width) and iterations (depth), the authors find that optimal budget allocation depends on the strategy, task, and total budget, and that strategy rankings shift with different budgets. They propose a measurement protocol that maps the seeds‑by‑iterations frontier and offers practical guidance for researchers.

arXiv AI
Aug 26

Exploit More, Explore Smarter for Budget-Constrained Agentic Search

The paper introduces ExTS, a tree‑search policy designed for budget‑constrained agentic search where evaluation and generation costs are high. ExTS treats expansion as a value‑of‑information decision, combining discriminative reward shaping, a stochastic virtual child, and quality‑conditioned branching to allocate budget more effectively. Experiments on prompt optimization, code generation, molecular structure elucidation, and agentic workflow optimization show ExTS matching or surpassing task‑specific baselines with an average gain of +5.5% using a single configuration, and the authors also present pilot‑run diagnostics to guide adaptation to different problem structures.

By Haoyang Fang, Bernie Wang
Hugging Face Trending Papers
Jul 30

SCOPE: Synthetic Conditional Objectives for Policy Evolution in Black-Box Combinatorial Optimization

Black-box combinatorial optimization requires systematically identifying high-quality solutions under a limited evaluation budget, yet the unknown objective function provides little guidance for deciding where the search should explore next. We introduce SCOPE, a general framework for Synthetic Conditional Objectives for Policy Evolution in Black-Box Combinatorial Optimization.

arXiv AI
Aug 20

Test-Time Scaling in the Wild: Why Exploitation, Not Exploration, Is the Bottleneck

The paper examines test‑time scaling (TTS) methods that use extra inference compute to improve language model outputs. Across five open‑ended benchmarks—medicine, law, finance, general chat, and creative writing—the study finds that increasing exploration (generating more candidates) consistently yields better top candidates, but exploitation (selecting the final output) remains weak due to poor reward‑model correlation. Only the Fusion approach, which synthesizes candidates, reliably improves results, yet it recovers only about 40% of the potential quality, indicating that the bottleneck lies in choosing from the candidate pool rather than generating it.

By Davide Romano, Kanak Raj, Jerrod Parker, Daniele Giofr\`e
arXiv Machine Learning
5d ago

EvoRank: LLM-Guided Evolution of Multi-Objective Learning-to-Rank Pipelines

EvoRank is an open autonomous ranking engineer that uses an LLM-guided evolutionary loop to automatically design complete Learning-to-Rank pipelines—including features, models, losses, and ensembles—for multi-objective e-commerce search. On the Expedia ICDM 2013 dataset, EvoRank converged within 50 iterations on interpretable pipelines that outperform an Optuna-tuned LambdaMART on 60k held-out queries and rank in the top 6 % of the original competition. The authors also introduce a headroom gate that predicts whether the evolutionary loop will be worthwhile before any LLM computation, and they release the system, auditing tools, and a catalog of failure modes to help teams apply the method to their own ranking stacks.

By Rayhan Patel, Shabaz Patel