Hugging Face Trending Papers

Budget-Aware LLM Discovery via Cost-Calibrated Frontier Utility

Large language models increasingly support scientific and algorithmic discovery through inference-time search over evaluated candidates. Existing adaptive discovery controllers assign credit based only on score progress, even though prompt length, retries, and guidance calls cause search actions to incur different token costs.

arXiv AI
Sep 2

One Policy, Any Budget: Internalizing Budget-Aware Search via Reinforcement Learning

The paper introduces AnySearch, a reinforcement‑learning framework that trains a single policy to perform budget‑aware search for large language models under any budget constraint. The training proceeds in two phases: first, the agent learns with explicit budget state injection and structured reasoning prompts under linearly decaying budgets; second, the scaffold is removed and the agent adapts to randomly sampled budgets that match deployment conditions. The reward combines answer accuracy and budget efficiency, with adaptive weighting to emphasize efficiency for high‑accuracy queries and reduce it for low‑accuracy ones. Experiments on seven QA benchmarks demonstrate that AnySearch outperforms baselines across all budget scales, generalizes to unseen constraints, and improves tool productivity without excessive token overhead.

By Xiaowei Sun, Jin Li, Yili Hong, Yikun Fu, Yanghua Xiao
arXiv AI
Aug 26

Exploit More, Explore Smarter for Budget-Constrained Agentic Search

The paper introduces ExTS, a tree‑search policy designed for budget‑constrained agentic search where evaluation and generation costs are high. ExTS treats expansion as a value‑of‑information decision, combining discriminative reward shaping, a stochastic virtual child, and quality‑conditioned branching to allocate budget more effectively. Experiments on prompt optimization, code generation, molecular structure elucidation, and agentic workflow optimization show ExTS matching or surpassing task‑specific baselines with an average gain of +5.5% using a single configuration, and the authors also present pilot‑run diagnostics to guide adaptation to different problem structures.

By Haoyang Fang, Bernie Wang
arXiv Machine Learning
Jun 3

Resource-Constrained Adaptive Inference for Sequential Pricing

arXiv:2606. 03736v1 Announce Type: cross Abstract: Resource-constrained pricing controllers can make fixed-price inference impossible: the controller's resource state may remove the target price neighborhood from the feasible set, even when every realized action has a known positive density.

By Ruicheng Ao, Jiashuo Jiang, David Simchi-Levi
arXiv Machine Learning
Jul 16

Where Should RL Post-Training Compute Go? Model Size, Search, Learning, and Feedback

arXiv:2607. 13389v1 Announce Type: new Abstract: Reinforcement Learning (RL) post-training is increasingly used to adapt foundation models for reasoning, planning, and feedback-driven robot-learning pipelines, but constrained post-training resources are often summarized by a single total FLOP budget.

By Patrick Wilhelm, Odej Kao