arXiv Machine Learning

Hill Sampling for Test-Time Scaling: A Simple and Better Alternative to Repeated Sampling, Evolution, and Training

arXiv Machine Learning
Sep 18

Sample Count Is Not Enough: Candidate-Generation Strategy Shapes the Energy and Performance of LLM Test-Time Scaling

The paper demonstrates that the number of candidates generated during test-time scaling of large language models does not fully capture the system cost. By comparing different generation schedules (e.g., one batched call versus multiple serial calls) while keeping the total candidate count fixed, the authors show that serial calls consume significantly more GPU energy and latency. The study suggests that reporting candidate count alone is insufficient; evaluations should also include generation schedule and GPU-level metrics.

By Mobina Kashaniyan, Ali Jannesari
arXiv Machine Learning
Jul 10

HeaPA: Difficulty-Aware Heap Sampling and On-Policy Query Augmentation for LLM Reinforcement Learning

arXiv:2601. 22448v2 Announce Type: replace Abstract: RLVR has become a standard recipe for training LLMs on reasoning tasks with verifiable outcomes, but when rollout generation dominates the cost, efficiency hinges on which prompts are sampled and when.

By Weiqi Wang, Xin Liu, Binxuan Huang, Hejie Cui, Rongzhi Zhang, Changlong Yu, Shuowei Jin, Jingfeng Yang, Qingyu Yin, Zhengyang Wang, Zheng Li, Yifan Gao, Priyanka Nigam, Bing Yin, Lihong Li, Yangqiu Song
arXiv Machine Learning
Aug 19

Agentic ESOpt: Fine-Tuning Long-Horizon LLM Agents with Minimal GPU Requirements

Agentic ESOpt proposes using evolution strategies (ES) instead of reinforcement learning to fine‑tune large language‑model agents for long‑horizon tasks. ES offers model scalability, flexibility, and better long‑horizon credit assignment, enabling full‑parameter optimization with minimal GPU memory. The framework samples parameter perturbations, evaluates agents with rewards, and updates online, achieving notable performance gains on WebArena‑Lite and in test‑time prompt‑parameter co‑evolution.

By Zhi Zheng, Rongsheng Chen, Yunpeng Ba, Zhenkun Wang, Yee Whye Teh, Wee Sun Lee
arXiv AI
Sep 7

LLM-Guided Program Evolution for Circle Packing: Breaking 10 Packomania Records for $28

The paper introduces Discovery Loop, a lightweight system that employs a large language model (LLM) to iteratively evolve optimization algorithms for the Packomania circle‑packing benchmark. Starting from a simple seed solver, the LLM proposes algorithmic improvements guided by a scoreboard and a history of prior ideas, evaluates each candidate against an independent verifier, and retains only successful changes. Within 15 iterations and a total LLM cost of $27.72, the system broke 10 Packomania records for N between 101 and 114, improving the best known solutions by 2.4%–5.4%. The work demonstrates that a cost‑efficient, LLM‑driven approach can rapidly advance state‑of‑the‑art solutions in a complex optimization domain, suggesting broader potential for democratizing automated scientific discovery.

By Wes Sander
arXiv AI
Sep 11

Why Sample What You Can Enumerate? Exact Policy Optimization for Genomic Tool Selection

The paper critiques the common reinforcement‑learning approach of sampling tool subsets when the full set of tools is enumerable, showing that sampling leads to degraded policy estimates and increased reward sparsity in genomic reasoning tasks. It proposes Full‑Group Policy Optimization (FGPO), which evaluates every tool subset and precomputes rewards in a table, thereby eliminating the need for frozen‑reasoner calls during training. Experiments across five frozen reasoners and three genomic benchmarks demonstrate that FGPO consistently outperforms GRPO, improving average scores by 6.75 points and reducing the number of invoked tools per question.

By Haoyue Liu, Xiaoyu Ma, Ye Chen, Zhichao Wang, Xiaoying Tang