arXiv Machine Learning By Yansen Zhang, Yilu Liu, Tianyu Liu, Jiamin Chen, Xiaokun Zhang, Kai Xie, Xue Liu, Chen Ma, Yiyan Qi

Budget-Aware LLM Discovery via Cost-Calibrated Frontier Utility

Read the original on arXiv Machine Learning →

arXiv:2607. 26828v1 Announce Type: new Abstract: Large language models increasingly support scientific and algorithmic discovery through inference-time search over evaluated candidates.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv AI
Sep 2

One Policy, Any Budget: Internalizing Budget-Aware Search via Reinforcement Learning

The paper introduces AnySearch, a reinforcement‑learning framework that trains a single policy to perform budget‑aware search for large language models under any budget constraint. The training proceeds in two phases: first, the agent learns with explicit budget state injection and structured reasoning prompts under linearly decaying budgets; second, the scaffold is removed and the agent adapts to randomly sampled budgets that match deployment conditions. The reward combines answer accuracy and budget efficiency, with adaptive weighting to emphasize efficiency for high‑accuracy queries and reduce it for low‑accuracy ones. Experiments on seven QA benchmarks demonstrate that AnySearch outperforms baselines across all budget scales, generalizes to unseen constraints, and improves tool productivity without excessive token overhead.

By Xiaowei Sun, Jin Li, Yili Hong, Yikun Fu, Yanghua Xiao
arXiv AI
Aug 26

Exploit More, Explore Smarter for Budget-Constrained Agentic Search

The paper introduces ExTS, a tree‑search policy designed for budget‑constrained agentic search where evaluation and generation costs are high. ExTS treats expansion as a value‑of‑information decision, combining discriminative reward shaping, a stochastic virtual child, and quality‑conditioned branching to allocate budget more effectively. Experiments on prompt optimization, code generation, molecular structure elucidation, and agentic workflow optimization show ExTS matching or surpassing task‑specific baselines with an average gain of +5.5% using a single configuration, and the authors also present pilot‑run diagnostics to guide adaptation to different problem structures.

By Haoyang Fang, Bernie Wang