Dynamic Budget Allocation for LLM Evaluation under Hard Resource Constraints
Read the original on arXiv Machine Learning →The paper introduces HARP, a hard-budget allocation method for evaluating large language models (LLMs) in multi-turn interactions where the time-to-event is partially observed due to resource limits. HARP guarantees that the total computational budget is never exceeded, reallocates unused budget, and provides lower predictive bounds (LPBs) with finite-sample coverage and unbiased metric estimates. Experiments on tasks such as jailbreaks, toxic content, and hallucinations demonstrate that HARP achieves near-nominal coverage with low variance while respecting the fixed budget.
Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.