arXiv Computer Vision

VisionMX: Unlocking Microscaling Post-Training Quantization for Vision Models

VisionMX introduces a post‑training microscaling (MX) quantization technique for vision models, addressing three key error sources identified in direct conversion: block‑scale representation, misalignment of small convolutional weight tensors, and underuse of signed codes for nonnegative activations. The method optimizes bounded weight rounding and incorporates a foldable affine correction for activations, improving performance across image classification, object detection, semantic segmentation, and low‑light image enhancement tasks. VisionMX outperforms direct conversion and other post‑training quantization baselines, especially in architectures most sensitive to MX conversion.

arXiv Computer Vision
Sep 18

MiX: Micro-Inverted-Scaling for End-to-End Low-Bit Vision-Language Model Acceleration

The paper introduces MiX, a micro‑inverted‑scaling format that replaces shared exponents with shared mantissas to avoid microscaling collapse in low‑bit vision‑language models. An adaptive dual‑format inference framework (MiX‑MX) maps this format to a custom accelerator, replacing multipliers with shifters. Experiments show 4.5‑bit MiX matches or outperforms NVFP4 accuracy while improving area efficiency by 25 % and delivering 2.3–4.5× speedup with 1.4–2.9× energy savings over the Focus accelerator.

By Yuan Liao, Jae-sun Seo
Hugging Face Trending Papers
Sep 17

MiX: Micro-Inverted-Scaling for End-to-End Low-Bit Vision-Language Model Acceleration

MiX: Micro-Inverted-Scaling for End-to-End Low-Bit Vision-Language Model Acceleration proposes a new quantization format that inverts the traditional microscaling approach by assigning private exponents to each element and a shared mantissa. The adaptive dual-format MiX-MX inference framework maps this format to a custom accelerator, replacing multipliers with shifters. Evaluations show that 4.5-bit MiX matches or surpasses NVFP4 accuracy on multimodal benchmarks while improving area efficiency by 25% and delivering 2.3–4.5× speedup with 1.4–2.9× energy reduction compared to the Focus accelerator.

arXiv Computation and Language
Aug 31

H-Scale: Hessian-Guided Scale Refinement for NVFP4 Sub-Byte LLM Inference

The paper introduces H-Scale, a lightweight post-processing technique for refining per-group scaling factors in NVFP4 quantized large language models. By using a diagonal second-order proxy from calibration activations, H-Scale selects hardware-valid scales that directly target layer output perturbation rather than just weight reconstruction error. Experiments on mainstream LLMs show that H-Scale improves NVFP4 baselines and brings several variants closer to BF16 performance without adding inference overhead.

By Hao Yu, Zheng Li, Dayiheng Liu, Jianwei Zhang