← Back to all news
arXiv AI October 2, 2026 By Chao Li, Shigeng Wang, Anbang Yao

The Devil Is in the Reconstruction Loss Scale: Rethinking Optimization in LLM Quantization

Read the original on arXiv AI →

The Flow has not summarised this story yet — read it at arXiv AI.

  • llms
  • efficiency
  • benchmarks

One email a morning, machine-written

One email a day, machine-written, one click to leave. We never share your address.

Related stories

arXiv Machine Learning
Aug 17

QUASAR: Lowering the Loss Floor of Quantization-Aware Training with Loss-Aware Reconstruction

arXiv:2608. 13966v1 Announce Type: new Abstract: As large language model inference shifts toward lower precision, post-training quantization (PTQ) becomes increasingly brittle, making quantization-aware training (QAT) essential for preserving model quality.

By Vincent Counathe, Ben Athiwaratkun, Christopher De Sa, Tianyi Zhang
llmsefficiency
More like this →
Hugging Face Trending Papers
Aug 14

QUASAR: Lowering the Loss Floor of Quantization-Aware Training with Loss-Aware Reconstruction

As large language model inference shifts toward lower precision, post-training quantization (PTQ) becomes increasingly brittle, making quantization-aware training (QAT) essential for preserving model...

llmsefficiency
More like this →
arXiv Machine Learning
Aug 17

Post-training Quantization for Hybrid Iterative Generative Models

arXiv:2608. 13932v1 Announce Type: new Abstract: Iterative Generative Models (IGMs) span autoregressive and diffusion paradigms, and hybrid variants that couple them can achieve remarkable image-generation fidelity.

By Jing Gao, Junyi Wu, Wei Wang, Yan Yan, Yao Zhao
diffusionefficiency
More like this →
arXiv AI
Jun 4

QuBLAST: A Framework for Quantizing Large Language Models with Block-Level Compression Approach and Activation Scaling Strategy

arXiv:2606. 04620v1 Announce Type: cross Abstract: LLMs have become the state-of-the-art algorithms for solving NLP tasks.

By Pasindu Wickramasinghe, Achyuta Muthuvelan, Rachmad Vidya Wicaksana Putra, Minghao Shao, Muhammad Shafique
llmsnlpefficiencybenchmarks
More like this →
arXiv AI
Jun 9

ScaleSweep: Accurate NVFP4 Post-Training Quantization of LLMs via Block Scale Initialization

arXiv:2606. 07618v1 Announce Type: cross Abstract: NVFP4 is a recently introduced hardware-supported FP4 format that improves the fidelity of 4-bit quantization through fine-grained block scales.

By Li Lin, Xiaojun Wan
llmsefficiency
More like this →
arXiv AI
Jun 18

HeRo-Q: A General Framework for Stable Low Bit Quantization via Hessian Conditioning

arXiv:2601. 21626v2 Announce Type: replace-cross Abstract: Post Training Quantization (PTQ), a mainstream model compression technique, often leads to the paradoxical 'low error, high loss' phenomenon because it focuses solely on minimizing quantization error.

By Jinhao Zhang, Yunquan Zhang, Zicheng yan, Boyang Zhang, Jun Sun, Daning Cheng
llmsefficiency
More like this →
About Pricing API Newsletter Sources Privacy Terms Refunds Accessibility Provider info Contact RSS

The Flow links to publishers and never republishes their articles. Summaries are machine-generated.

v1.1.0 · 5f852ea