arXiv AI By Rodrigo Guedes de Souza, Alison R. Panisson

Who Thinks Best Depends on How Long You Let Them: Budget-Dependent Rankings in LLM Evaluation

Read the original on arXiv AI →

arXiv:2608. 12150v1 Announce Type: new Abstract: Standard evaluation of large language models assumes stable model rankings across inference conditions.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Aug 20

Can a Lightweight Multimodal Model Estimate LLM Reasoning Performance? A Study for Compute-Optimal Document Inference

The paper introduces BudgetDoc, a multimodal benchmark that explicitly supervises the trade‑off between inference budget and performance across three document tasks. Using this benchmark, the authors train DRB, a lightweight 1‑billion‑parameter pre‑flight estimator (SigLIP‑2 + Qwen3‑0.6B) that predicts ordinal model performance for different budget levels and achieves a weighted F1 of 0.753. When DRB dynamically allocates reasoning budgets to five frontier models on three datasets, it matches or improves F1 scores compared to always‑maximum‑budget baselines in 9 of 15 configurations while dramatically cutting cost, and preliminary tests suggest it may generalize to cross‑model selection.

By Zishan Ahmad, Vishal Vaddina