arXiv Machine Learning

Online Verification of Language Model Responses Under Cost Constraints

The paper introduces OMVV, an online multi-verifier algorithm that maintains a pool of weak verifiers with varying costs and performance. It adaptively selects a verifier each round using an online score combiner and exponential-weights routing, providing distribution-free guarantees on false-accept and false-reject rates. Experiments on reasoning benchmarks show OMVV achieves higher accuracy at lower verification cost than any single fixed verifier across different budgets.

arXiv Computation and Language
Aug 28

TRACES: Tagging Reasoning Steps for Adaptive Cost-Efficient Early-Stopping

TRACES (Tagging Reasoning Steps for Adaptive Cost‑Efficient Early‑Stopping) is a lightweight framework that tags reasoning steps of large‑language models in real time, enabling adaptive, cost‑efficient early stopping during inference. By monitoring the types of steps generated, the method identifies when models shift their reasoning after arriving at a correct answer, allowing for interpretable stopping criteria. Experiments on mathematical reasoning benchmarks (MATH500, GSM8K, AIME) and knowledge benchmarks (MMLU, GPQA) show token reductions of 20–50% while preserving accuracy, with more conservative thresholds needed for harder tasks such as BeyondAIME and IMO AnswerBench.

By Yannis Belkhiter, Seshu Tirupathi, Giulio Zizzo, John D. Kelleher
arXiv Machine Learning
Sep 2

Online Self-Weighted Fine-Tuning

Online Self-Weighted Fine‑Tuning (OSW‑FT) augments standard supervised fine‑tuning by adding online, trajectory‑level weighting: for each query the model estimates its current success rate from a small number of inference‑only rollouts and rescales the SFT loss accordingly. The method keeps the optimization direction anchored to the expert trajectory while adapting the update magnitude online, and it is shown to be unbiased for any finite rollout count with a convergence analysis. Across Qwen3 models from 0.6B to 4B, OSW‑FT consistently outperforms plain SFT on challenging benchmarks such as AIME, achieving a favorable compute‑performance trade‑off with only two online rollouts.

By Haiquan Wen, Yiwei He, Bei Peng, Guangliang Cheng
arXiv Machine Learning
3d ago

Which LLM to pick? Online Active Model Selection for Large Language Models

The paper introduces ONLINE LLM PICKER, a framework for active model selection of large language models in streaming settings. It selects the most informative prompts for annotation within a limited budget, enabling the identification of the best or near‑best model among many candidates. Experiments on 10 datasets and over 130 language models show up to 71.67% savings in annotation cost and a reduction in regret by up to 2.51× when using the chosen model for sequential generation.

By Alessandro Turrin, Patrik Okanovic, Torsten Hoefler, Nezihe Merve G\"urel