The paper introduces BudgetDoc, a multimodal benchmark that explicitly supervises the trade‑off between inference budget and performance across three document tasks. Using this benchmark, the authors train DRB, a lightweight 1‑billion‑parameter pre‑flight estimator (SigLIP‑2 + Qwen3‑0.6B) that predicts ordinal model performance for different budget levels and achieves a weighted F1 of 0.753. When DRB dynamically allocates reasoning budgets to five frontier models on three datasets, it matches or improves F1 scores compared to always‑maximum‑budget baselines in 9 of 15 configurations while dramatically cutting cost, and preliminary tests suggest it may generalize to cross‑model selection.
By Zishan Ahmad, Vishal Vaddina
arXiv:2606. 25986v1 Announce Type: new Abstract: We study whether a scaling-law-style inference-compute frontier appears in limit order book prediction.
By C. Evans Hedges
arXiv:2605. 06485v2 Announce Type: replace-cross Abstract: Large language models (LLMs) have transformed artificial intelligence, but their computational requirements remain prohibitive for most users.
By Nii Osae Osae Dade, Tony Morri, Moinul Hossain Rahat, Sayandip Pal, Rickston Pinto
SpecQuant is a training‑free framework that merges speculative decoding with multi‑parent quantization to enable adaptive, efficient inference of large language models. It generates several quantized variants (INT4, FP8, FP16) from a single base model and routes queries to the appropriate variant based on predicted complexity, using lightweight models for simple tasks and full‑precision models for complex reasoning. Evaluations on Qwen2.5 models across MMLU, AlpacaEval, and GSM8K show 35‑43% speedups with less than 2% accuracy loss, facilitating practical on‑device LLM deployment without specialized infrastructure.
By Harish KB, Jagadeeswaran M, Pradheep P, Yuvanesh S, Sivakumar T
arXiv:2607. 04244v1 Announce Type: new Abstract: This report describes our approach to the Efficient Qwen Competition, where the goal is to enable low-latency serving of Qwen3.
By Jaeyeon Kim, Jewon Lee, Bo-Kyeong Kim
Frontier large language models (LLMs) are examined as batch optimizers in both continuous and discrete settings. The study finds that while LLMs perform competitively in zero‑shot optimization of numerical test functions, their performance is less robust than classical non‑LLM methods. However, LLMs excel in semantically rich, discrete spaces that resemble their pretraining data, demonstrating strong batch optimization behavior in such contexts.
By Frank Hu, Shriram Chennakesavalu, David Graff
arXiv:2608. 08020v1 Announce Type: new Abstract: Test-time compute scaling is a primary driver of performance in large reasoning models (LRMs), but extreme inefficiency bounds current approaches, shifting the critical question from \emph{how much} compute to spend, to \emph{where} to allocate it.
By Lijie Yang, Hongyin Luo, Tri Dao, Ravi Netravali
arXiv:2607. 17733v1 Announce Type: cross Abstract: 4-bit quantization enables efficient LLM inference, but suffers from significant accuracy degradation due to outliers.
By Simla Burcu Harma, Danila Mishin, Zhengyuan Su, Ayan Chakraborty, Elizaveta Kostenok, Dongho Ha, Babak Falsafi, Martin Jaggi, Yunho Oh, Amir Yazdanbakhsh
arXiv:2607. 14557v1 Announce Type: new Abstract: Diffusion Multimodal Large Language Models (DMLLMs) are highly effective for multimodal reasoning, yet their inference efficiency is significantly hindered by fixed-length generation constraints.
By Qicheng Zhao, Qi Sun, Zheyu Yan
arXiv:2608. 10523v1 Announce Type: cross Abstract: \texttt{TensorSketch} by~\cite{pham2013fast,kar2012random} provides efficient sketching algorithms for high-dimensional polynomial kernels $\vec{x}^{\otimes p} \in \R^{d^p}$.
By Amit Sharma, Mohammad Azhar Khan, Rameshwar Pratap, Keegan Kang