arXiv:2606. 25432v1 Announce Type: new Abstract: Inference efficiency is typically pursued by shrinking the model: distillation, pruning, quantization, and sparse routing each lower per-token cost while treating token count as fixed.
By DatologyAI, :, Matthew L. Leavitt, Siddharth Joshi, Haoli Yin, Rishabh Adiga, Haakon Mongstad, Alvin Deng, David Schwab, Bogdan Gaza, Ari Morcos
arXiv:2608. 09351v1 Announce Type: cross Abstract: Test-time scaling improves LLM accuracy but multiplies inference cost, making the accuracy gained per unit of compute the metric that matters in deployment.
By Nikita Kozodoi, Zainab Afolabi, Jack Butler
arXiv:2609.25809v1 Announce Type: new
Abstract: Fine-grained mixture-of-experts (MoE) architectures have become a mainstream design for open-weight LLMs, with hundreds of experts and increasingly man...
By Yuanteng Chen, Qiwei Lai, Chen Tianqi, Peisong Wang, Yuantian Shao, Nanxin Zeng, Zhilei Liu, Chuangyi Li, Jing Liu, Jian Cheng
Parameter-efficient fine-tuning (PEFT) and low-bit quantization are now standard tools for adapting language models under tight compute budgets, yet their interaction is most often studied on billion-parameter models where the design space is expensive to explore. We ask a complementary question: on a specific, fully reproducible 60M-parameter encoder-decoder model (T5-small) and a single-table text-to-SQL benchmark (WikiSQL), how much task accuracy does each efficiency knob actually cost?
OBC‑Prune introduces an outcome‑based calibration approach for pruning large reasoning models, focusing on the causal importance of each reasoning sentence rather than uniform activation salience. By pairing correct and incorrect rollouts and using intervention‑based analysis, it assigns per‑token weights that guide one‑shot pruning methods such as SparseGPT, Wanda, and ALPS. Experiments on DeepSeek‑R1‑Distill‑Qwen models show consistent accuracy gains and shorter reasoning traces across multiple benchmarks at 40‑50% sparsity.
By Ha Lan Nguyen, Huy Hoang Tran, Trac-Duy Tran, Dung D. Le
arXiv:2606. 18557v1 Announce Type: new Abstract: A rule-based logic solver resolves every instance in our benchmark in under 50 microseconds with 100% accuracy; the best frontier language model reaches 65% at best and drops to 23.
By Patrick Cooper, Alvaro Velasquez