arXiv Machine Learning By Shai Feldman, Yaniv Romano

Dynamic Budget Allocation for LLM Evaluation under Hard Resource Constraints

Read the original on arXiv Machine Learning →

The paper introduces HARP, a hard-budget allocation method for evaluating large language models (LLMs) in multi-turn interactions where the time-to-event is partially observed due to resource limits. HARP guarantees that the total computational budget is never exceeded, reallocates unused budget, and provides lower predictive bounds (LPBs) with finite-sample coverage and unbiased metric estimates. Experiments on tasks such as jailbreaks, toxic content, and hallucinations demonstrate that HARP achieves near-nominal coverage with low variance while respecting the fixed budget.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv Computation and Language
Sep 15

When Agents Slow Down: Understanding LLM Agents' Test-Time Strategies via Elo-per-token Analysis

arXiv:2609.15309v1 Announce Type: new Abstract: Large language model (LLM) agents allocate test-time compute adaptively as they revise solutions, use tools, explore alternatives, and decide when to s...

By Kaiyuan Liu, Qiuyang Mang, Bo Peng, Wenhao Chai, Hanchen Li, Shreyas Pimpalgaonkar, Luke Zettlemoyer, Alex Dimakis, Alvin Cheung