arXiv:2606. 04115v1 Announce Type: cross Abstract: Quantizing large language models (LLMs) to low-precision floating-point representations is central to efficient deployment, yet applying a single bit-width uniformly across all layers is sub-optimal in terms of both performance and accuracy.
By Giuseppe Franco, Ian Colbert, Pablo Monteagudo-Lago, Felix Marty, Nicholas Fraser
arXiv:2606. 06527v1 Announce Type: cross Abstract: Energy-efficient edge inference requires reducing arithmetic cost, memory traffic, and hardware overhead.
By Ovishake Sen, Venkata Nithin Kamineni, Daniel Lobo, Swarup Bhunia, Rickard Ewetz, Baibhab Chatterjee
The paper introduces MiX, a micro‑inverted‑scaling format that replaces shared exponents with shared mantissas to avoid microscaling collapse in low‑bit vision‑language models. An adaptive dual‑format inference framework (MiX‑MX) maps this format to a custom accelerator, replacing multipliers with shifters. Experiments show 4.5‑bit MiX matches or outperforms NVFP4 accuracy while improving area efficiency by 25 % and delivering 2.3–4.5× speedup with 1.4–2.9× energy savings over the Focus accelerator.
By Yuan Liao, Jae-sun Seo
Squeeze10-LLM is a staged mixed‑precision post‑training quantization framework that reduces 16‑bit LLM weights to an average of 1.6 bits per weight by assigning 80% of weights to 1 bit and 20% to 4 bits. It introduces Post‑Binarization Activation Robustness (PBAR), a weight significance metric that considers activation impact, and Full Information Activation Supervision (FIAS), a strategy that preserves activation information to limit error propagation. Experiments on LLaMA and LLaMA2 demonstrate that Squeeze10‑LLM achieves state‑of‑the‑art performance for sub‑2‑bit weight‑only quantization, raising average accuracy from 43% to 56% on six zero‑shot classification tasks.
By Qingcheng Zhu, Yangyang Ren, Linlin Yang, Yanjing Li, Sheng Xu, Haodong Zhu, Juan Zhang, Runqi Wang, Baochang Zhang
FAME is an FPGA-based platform that evaluates approximate multipliers directly in hardware, eliminating slow CPU/GPU LUT emulation and reducing evaluation time for DNN inference. It also introduces a pattern-guided retraining method that uses multiplier-specific patterns to recover accuracy losses. Experiments on ResNet‑18 and MobileNetV2 over ImageNet show up to 3.47× faster multiplier evaluation and a 65.5% accuracy improvement over prior retraining approaches.
By Rappy Saha, Nima Amirafshar, Jude Haris, Nima Taherinejad, Jos\'e Cano
arXiv:2607. 28418v1 Announce Type: cross Abstract: Pruning is a promising approach for improving the efficiency of LLMs.
By Haozhe Hu, Hao Wu, Peiran Yin, Chao Han, Yunpu Ma, Xiaoyu Shen
MiX: Micro-Inverted-Scaling for End-to-End Low-Bit Vision-Language Model Acceleration proposes a new quantization format that inverts the traditional microscaling approach by assigning private exponents to each element and a shared mantissa. The adaptive dual-format MiX-MX inference framework maps this format to a custom accelerator, replacing multipliers with shifters. Evaluations show that 4.5-bit MiX matches or surpasses NVFP4 accuracy on multimodal benchmarks while improving area efficiency by 25% and delivering 2.3–4.5× speedup with 1.4–2.9× energy reduction compared to the Focus accelerator.
arXiv:2601. 07475v2 Announce Type: replace-cross Abstract: The emergence of fine-grained numerical formats like NVFP4 presents new opportunities for efficient Large Language Model (LLM) inference.
By Haoqian Meng, Yilun Luo, Yafei Zhao, Wenyuan Liu, Peng Zhang, Xindian Ma
arXiv:2606. 06527v2 Announce Type: replace-cross Abstract: Energy-efficient neural-network inference at the edge requires reducing arithmetic cost, memory traffic, computation energy, and storage overhead while maintaining acceptable accuracy.
By Ovishake Sen, Venkata Nithin Kamineni, Daniel Lobo, Swarup Bhunia, Rickard Ewetz, Baibhab Chatterjee
arXiv:2608. 15602v1 Announce Type: cross Abstract: While binary quantization theoretically promises extreme compression and acceleration for Large Language Models (LLMs), existing research often overlooks the necessity of specialized hardware kernels, thus failing to unleash the full acceleration potential due to persistent reliance on expensive floating-point arithmetic or runtime dequantization overheads.
By Qingyao Yang, Runming Yang, He Xiao, Wendong Xu, Junyu Chen, Haobo Liu, Chenchen Ding, Ruihan Hu, Yik-Chung Wu, Ngai Wong
arXiv:2607. 17733v1 Announce Type: cross Abstract: 4-bit quantization enables efficient LLM inference, but suffers from significant accuracy degradation due to outliers.
By Simla Burcu Harma, Danila Mishin, Zhengyuan Su, Ayan Chakraborty, Elizaveta Kostenok, Dongho Ha, Babak Falsafi, Martin Jaggi, Yunho Oh, Amir Yazdanbakhsh
arXiv:2606. 17249v1 Announce Type: cross Abstract: The dominant trajectory of modern machine learning has been to scale up: larger models, larger accelerators, larger memory budgets.
By Emre Can Kizilates