The paper introduces RAMP, a method for robust adaptive mixed‑precision quantization of vision models on edge CPUs. It evaluates 13 sensitivity metrics across four neural networks, finding that Jensen‑Shannon Divergence consistently identifies layers that can be safely quantized. Using K‑Means clustering on these metrics, RAMP achieves near‑lossless accuracy with an average 1.81× speed‑up, while cautioning against excluding low‑speed‑up layers that can fragment the computational graph.
By David Poblaci\'on-Criado, Dario Garcia-Gasulla, Eduardo Quinones
The paper introduces MiX, a micro‑inverted‑scaling format that replaces shared exponents with shared mantissas to avoid microscaling collapse in low‑bit vision‑language models. An adaptive dual‑format inference framework (MiX‑MX) maps this format to a custom accelerator, replacing multipliers with shifters. Experiments show 4.5‑bit MiX matches or outperforms NVFP4 accuracy while improving area efficiency by 25 % and delivering 2.3–4.5× speedup with 1.4–2.9× energy savings over the Focus accelerator.
By Yuan Liao, Jae-sun Seo
MiX: Micro-Inverted-Scaling for End-to-End Low-Bit Vision-Language Model Acceleration proposes a new quantization format that inverts the traditional microscaling approach by assigning private exponents to each element and a shared mantissa. The adaptive dual-format MiX-MX inference framework maps this format to a custom accelerator, replacing multipliers with shifters. Evaluations show that 4.5-bit MiX matches or surpasses NVFP4 accuracy on multimodal benchmarks while improving area efficiency by 25% and delivering 2.3–4.5× speedup with 1.4–2.9× energy reduction compared to the Focus accelerator.
arXiv:2607. 20981v1 Announce Type: new Abstract: Efficient multimodal inference is increasingly constrained not only by model quality or FLOP count, but also by the cost of preserving, moving, routing, caching, and quantizing multimodal representations under latency, memory, and energy constraints.
By Jay Gor, Karm Dave, Akshita Abrol, Rajesh Gupta, Sudeep Tanwar, Zhengkui Wang
Deploying deep learning models on edge CPUs is bottlenecked by computational and memory constraints. Mixed-precision quantization promises to reduce inference latency while preserving accuracy. Howeve...
The paper introduces Llama-Mobile, a framework that quantizes vision‑language models for efficient mobile deployment. It uses a quantization pipeline that generates training data from the model itself, eliminating the need for the original training setup, and employs a novel 2.7‑bit‑per‑parameter format optimized for Arm CPUs. Applying this method, the authors compress the Llama 3.2 11B Vision Instruct model to 3.7 GB with 8‑bit activations while maintaining strong performance on visual question answering tasks.
By Luka Ribar, Jeevan Bhoot, Douglas Orr