Hugging Face Blog

Up to 3.2x Faster Inference with LFM2.5-DSpark

arXiv AI
Aug 20

Can a Lightweight Multimodal Model Estimate LLM Reasoning Performance? A Study for Compute-Optimal Document Inference

The paper introduces BudgetDoc, a multimodal benchmark that explicitly supervises the trade‑off between inference budget and performance across three document tasks. Using this benchmark, the authors train DRB, a lightweight 1‑billion‑parameter pre‑flight estimator (SigLIP‑2 + Qwen3‑0.6B) that predicts ordinal model performance for different budget levels and achieves a weighted F1 of 0.753. When DRB dynamically allocates reasoning budgets to five frontier models on three datasets, it matches or improves F1 scores compared to always‑maximum‑budget baselines in 9 of 15 configurations while dramatically cutting cost, and preliminary tests suggest it may generalize to cross‑model selection.

By Zishan Ahmad, Vishal Vaddina
arXiv Machine Learning
Sep 21

SpecQuant: Speculative Decoding with Multi-Parent Quantization for Adaptive LLM Inference

SpecQuant is a training‑free framework that merges speculative decoding with multi‑parent quantization to enable adaptive, efficient inference of large language models. It generates several quantized variants (INT4, FP8, FP16) from a single base model and routes queries to the appropriate variant based on predicted complexity, using lightweight models for simple tasks and full‑precision models for complex reasoning. Evaluations on Qwen2.5 models across MMLU, AlpacaEval, and GSM8K show 35‑43% speedups with less than 2% accuracy loss, facilitating practical on‑device LLM deployment without specialized infrastructure.

By Harish KB, Jagadeeswaran M, Pradheep P, Yuvanesh S, Sivakumar T
arXiv Machine Learning
Sep 4

Frontier LLMs are effective batch optimizers: Assessing reasoning models in continuous and discrete settings

Frontier large language models (LLMs) are examined as batch optimizers in both continuous and discrete settings. The study finds that while LLMs perform competitively in zero‑shot optimization of numerical test functions, their performance is less robust than classical non‑LLM methods. However, LLMs excel in semantically rich, discrete spaces that resemble their pretraining data, demonstrating strong batch optimization behavior in such contexts.

By Frank Hu, Shriram Chennakesavalu, David Graff
arXiv AI
Aug 11

Thought-Level Beam Search for Reasoning

arXiv:2608. 08020v1 Announce Type: new Abstract: Test-time compute scaling is a primary driver of performance in large reasoning models (LRMs), but extreme inefficiency bounds current approaches, shifting the critical question from \emph{how much} compute to spend, to \emph{where} to allocate it.

By Lijie Yang, Hongyin Luo, Tri Dao, Ravi Netravali