Up to 3.2x Faster Inference with LFM2.5-DSpark
Read the original on Hugging Face Blog →The Flow has not summarised this story yet — read it at Hugging Face Blog.
The Flow has not summarised this story yet — read it at Hugging Face Blog.
The paper introduces BudgetDoc, a multimodal benchmark that explicitly supervises the trade‑off between inference budget and performance across three document tasks. Using this benchmark, the authors train DRB, a lightweight 1‑billion‑parameter pre‑flight estimator (SigLIP‑2 + Qwen3‑0.6B) that predicts ordinal model performance for different budget levels and achieves a weighted F1 of 0.753. When DRB dynamically allocates reasoning budgets to five frontier models on three datasets, it matches or improves F1 scores compared to always‑maximum‑budget baselines in 9 of 15 configurations while dramatically cutting cost, and preliminary tests suggest it may generalize to cross‑model selection.
arXiv:2606. 25986v1 Announce Type: new Abstract: We study whether a scaling-law-style inference-compute frontier appears in limit order book prediction.
arXiv:2605. 06485v2 Announce Type: replace-cross Abstract: Large language models (LLMs) have transformed artificial intelligence, but their computational requirements remain prohibitive for most users.
SpecQuant is a training‑free framework that merges speculative decoding with multi‑parent quantization to enable adaptive, efficient inference of large language models. It generates several quantized variants (INT4, FP8, FP16) from a single base model and routes queries to the appropriate variant based on predicted complexity, using lightweight models for simple tasks and full‑precision models for complex reasoning. Evaluations on Qwen2.5 models across MMLU, AlpacaEval, and GSM8K show 35‑43% speedups with less than 2% accuracy loss, facilitating practical on‑device LLM deployment without specialized infrastructure.