Knowing When to Stop: Bayesian Optimal Stopping for LLM Evaluations
arXiv:2608. 14425v1 Announce Type: new Abstract: LLM evaluations often use fixed sampling budgets, testing every item the same number of times even after estimates are precise.
The paper investigates how to decide when to stop collecting posterior samples in classification‑oriented adaptive sensing. It shows that a simple threshold‑based plug‑in rule does not guarantee the desired confidence level, and proposes calibrated fixed‑sample and finite‑horizon sequential stopping rules that control the false‑declaration probability. Experiments on MNIST demonstrate that the sequential rule can reduce sensing cost the most, and that a curtailment strategy can save up to 62% of posterior samples while maintaining accuracy.
arXiv:2608. 14425v1 Announce Type: new Abstract: LLM evaluations often use fixed sampling budgets, testing every item the same number of times even after estimates are precise.
arXiv:2602. 05395v2 Announce Type: replace-cross Abstract: A simple strategy for improving LLM accuracy, especially in math and reasoning problems, is to sample multiple responses and submit the answer most consistently reached.
The paper proposes a classification-oriented adaptive sensing method that uses posterior sampling from diffusion models. It leverages the closed-form posterior covariance of a class-conditional Gaussian mixture model to separate within-class and between-class uncertainty, estimating these terms from diffusion posterior samples via calibrated soft classifier outputs. Experiments on MNIST and CIFAR-10 demonstrate that this approach can achieve better classification accuracy for a given measurement cost compared to reconstruction-oriented methods, while also quantifying the associated reconstruction quality.
The paper studies how voting over multiple large language model (LLM) responses can be optimized under a fixed call budget. It introduces a recoverability threshold that quantifies the gap between discovering a correct answer and ensuring it wins the plurality vote, showing that the candidate set can only grow while the set of reachable winners can only shrink. The authors also present a gold‑free locking certificate that identifies the earliest prefix where all remaining continuations produce the same fixed‑budget output, and demonstrate empirical gains in accuracy and call efficiency through input permutation and exact locking.
arXiv:2601. 07094v2 Announce Type: replace-cross Abstract: Bayesian optimization (BO) iteratively fits a Gaussian process (GP) surrogate to accumulated evaluations and selects new queries via an acquisition function.
arXiv:2606. 15237v1 Announce Type: cross Abstract: Ensemble classifiers are predictive models that combine the results of simpler base models, often by majority vote.
arXiv:2608. 15520v1 Announce Type: new Abstract: A multimodal system may begin inference holding only some of its inputs and may acquire the rest at a cost.
arXiv:2605.18163v2 Announce Type: replace Abstract: Hallucination correction is not a one-direction problem. We show that intermediate layers are neither uniformly more truthful than final layers nor...
The paper introduces Cost-Aware Sequential Hypothesis Testing (CASHT), where a decision-maker selects sensing actions with varying random costs to identify the true hypothesis under an average-error constraint while minimizing expected total cost. For fixed costs, the optimal expected total cost scales as Θ(log(1/δ)) and can be achieved by Multihypothesis Sequential Probability Ratio Test-based procedures. The authors extend the framework to random costs under ex-post and ex-ante revelation models, analyze when action cancellation reduces cost, and demonstrate through simulations that CA variants consistently lower total cost compared to classical methods.
The paper introduces RouteCert, a method for ensuring risk control in multimodal systems that acquire inputs adaptively. It shows that conditional calibration can remain valid even when the acquisition policy determines the calibration group, and provides two finite‑sample constructions: threshold‑free routing with terminal‑pattern calibration and simultaneous validation of policy‑pattern pairs. Experiments on a clinical ECG task and masked multimodal benchmarks demonstrate that RouteCert achieves low disagreement rates and competitive answered fractions while validating each acquisition stage separately.
arXiv:2506. 24007v5 Announce Type: replace-cross Abstract: This study investigates minimax and Bayes optimal strategies for fixed-budget best-arm identification.
arXiv:2607. 12928v1 Announce Type: new Abstract: We study the online binary sequential calibration problem.