arXiv AI By Toby D. Pilditch

Knowing When to Stop: Bayesian Optimal Stopping for LLM Evaluations

Read the original on arXiv AI →

arXiv:2608. 14425v1 Announce Type: new Abstract: LLM evaluations often use fixed sampling budgets, testing every item the same number of times even after estimates are precise.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Computer Vision
Sep 21

How Many Posterior Samples? Calibrated Stopping for Adaptive Sensing

The paper investigates how to decide when to stop collecting posterior samples in classification‑oriented adaptive sensing. It shows that a simple threshold‑based plug‑in rule does not guarantee the desired confidence level, and proposes calibrated fixed‑sample and finite‑horizon sequential stopping rules that control the false‑declaration probability. Experiments on MNIST demonstrate that the sequential rule can reduce sensing cost the most, and that a curtailment strategy can save up to 62% of posterior samples while maintaining accuracy.

By Vincent Corlay, Andriy Enttsel
Hugging Face Trending Papers
Aug 18

Judge, Retrieve, or Abstain: Uncertainty-Guarded LLM Judging with Provable Risk Guarantees

The paper introduces a risk‑controlled framework for using large language models (LLMs) as judges in tasks without reference answers. By calibrating uncertainty thresholds on a held‑out set, the method ensures that the false discovery rate of accepted verdicts stays below a user‑specified level α with high probability, using finite‑sample Clopper–Pearson intervals. When the parametric judge lacks confidence, the instance is routed to a retrieval‑augmented mode with a second calibrated threshold, preserving the error guarantee while achieving higher coverage than single‑mode baselines.

arXiv AI
Jul 23

Statistical Early Stopping for Reasoning Models

arXiv:2602. 13935v2 Announce Type: replace Abstract: While LLMs have seen substantial improvement in reasoning capabilities, they also sometimes overthink, generating unnecessary reasoning steps, particularly under uncertainty, given ill-posed or ambiguous queries.

By Yangxinyu Xie, Tao Wang, Soham Mallick, Yan Sun, Georgy Noarov, Mengxin Yu, Tanwi Mallick, Weijie J. Su, Edgar Dobriban