MicroQonv: Reshaping Convolution Tensors for Efficient Microscaling in Training and Inference
Read the original on arXiv AI →The Flow has not summarised this story yet — read it at arXiv AI.
The Flow has not summarised this story yet — read it at arXiv AI.
arXiv:2607. 24953v1 Announce Type: cross Abstract: Reducing training precision is a key lever for improving the e ciency of large language model (LLM) training, but pushing beyond FP8 to 4-bit oating point (FP4) remains challenging due to instability during optimization.
The paper introduces RAMP, a method for robust adaptive mixed‑precision quantization of vision models on edge CPUs. It evaluates 13 sensitivity metrics across four neural networks, finding that Jensen‑Shannon Divergence consistently identifies layers that can be safely quantized. Using K‑Means clustering on these metrics, RAMP achieves near‑lossless accuracy with an average 1.81× speed‑up, while cautioning against excluding low‑speed‑up layers that can fragment the computational graph.
arXiv:2606. 04620v1 Announce Type: cross Abstract: LLMs have become the state-of-the-art algorithms for solving NLP tasks.
arXiv:2605. 25054v2 Announce Type: replace-cross Abstract: Deploying deep neural networks on resource-constrained 6G edge devices demands aggressive compression with minimal accuracy loss.
Deploying deep learning models on edge CPUs is bottlenecked by computational and memory constraints. Mixed-precision quantization promises to reduce inference latency while preserving accuracy. Howeve...
arXiv:2606. 07618v1 Announce Type: cross Abstract: NVFP4 is a recently introduced hardware-supported FP4 format that improves the fidelity of 4-bit quantization through fine-grained block scales.