arXiv:2606. 26836v1 Announce Type: new Abstract: Existing benchmarks typically report accuracy for a single model on a single run.
By Bradley Fowler, Ryan Smith, Daniel Thi Graviet, William Myers, Joshua Greaves, Narmeen Fatimah Oozeer, Ant\'ia Garc\'ia, Philip Quirke, Amirali Abdullah, Fazl Barez, Shriyash Kaustubh Upadhyay
arXiv:2606. 24083v1 Announce Type: cross Abstract: "Talk short.
By Morayo Danielle Adeyemi, Ryan A. Rossi, Franck Dernoncourt
The paper introduces BudgetDoc, a multimodal benchmark that explicitly supervises the trade‑off between inference budget and performance across three document tasks. Using this benchmark, the authors train DRB, a lightweight 1‑billion‑parameter pre‑flight estimator (SigLIP‑2 + Qwen3‑0.6B) that predicts ordinal model performance for different budget levels and achieves a weighted F1 of 0.753. When DRB dynamically allocates reasoning budgets to five frontier models on three datasets, it matches or improves F1 scores compared to always‑maximum‑budget baselines in 9 of 15 configurations while dramatically cutting cost, and preliminary tests suggest it may generalize to cross‑model selection.
By Zishan Ahmad, Vishal Vaddina
arXiv:2607. 08665v1 Announce Type: new Abstract: Routing among large language models (LLMs) trades response quality against serving cost, motivated by the reported gap between deployed routers and a per-instance oracle.
By Teng-Ruei Chen
arXiv:2609.13149v1 Announce Type: new
Abstract: For local large language model agents, active context is a scarce resource: memory capacity, prefill latency, cache growth, and service objectives all...
By Aditya Karnam Gururaj Rao, Arjun Jaggi
arXiv:2608. 13571v1 Announce Type: cross Abstract: When a language model fails to answer a query on the first attempt, an agentic system retries, consuming additional tokens each time.
By Heming Fu, Shan Lin, Qianqian Xie, Guojun Xiong