← Back to all news
Hugging Face Blog March 18, 2024

Quanto: a PyTorch quantization backend for Optimum

Read the original on Hugging Face Blog →

The Flow has not summarised this story yet — read it at Hugging Face Blog.

  • efficiency

One email a morning, machine-written

One email a day, machine-written, one click to leave. We never share your address.

Related stories

arXiv AI
Aug 13

VQ-bench: A Composable Vector Quantization Framework

arXiv:2608. 11240v1 Announce Type: new Abstract: Vector quantization is an old problem but has recently become central to AI infrastructure.

By Ashwin Padaki, Amir Ingber, Edo Liberty
efficiencybenchmarks
More like this →
arXiv Machine Learning
Jun 2

Inner Product Aware Quantization: Provably Fast, Accurate, and Adaptive Algorithms

arXiv:2606. 00289v1 Announce Type: new Abstract: Quantization is a fundamental tool used to compress datasets, neural network weights, and memory usage in a range of computational tasks.

By Nathan White, Krish Singal
efficiencybenchmarks
More like this →
arXiv Machine Learning
Jun 3

WaterSIC: Information-Theoretically (Near) Optimal Linear Layer Quantization

arXiv:2603. 04956v2 Announce Type: replace Abstract: This paper considers the problem of converting a given dense linear layer to low precision.

By Egor Lifar, Semyon Savkin, Or Ordentlich, Yury Polyanskiy
llmsefficiencybenchmarks
More like this →
arXiv Machine Learning
Jun 19

CAGE: Curvature-Aware Gradient Estimation For Accurate Quantization-Aware Training

arXiv:2510. 18784v3 Announce Type: replace Abstract: Despite significant work on low-bit quantization-aware training (QAT), there is still an accuracy gap between such techniques and native training.

By Soroush Tabesh, Mher Safaryan, Andrei Panferov, Alexandra Volkova, Dan Alistarh
llmsfine-tuningefficiencybenchmarks
More like this →
arXiv Machine Learning
Jul 10

KronQ: LLM Quantization via Kronecker-Factored Hessian

arXiv:2607. 07964v1 Announce Type: new Abstract: Post-training quantization (PTQ) is a widely adopted technique for compressing large language models (LLMs) without retraining.

By Donghyun Lee, Yuhang Li, Ruokai Yin, Priyadarshini Panda
llmsefficiency
More like this →
arXiv Machine Learning
Jun 16

QuantKAN: A Unified Quantization Framework for Kolmogorov Arnold Networks

arXiv:2511. 18689v3 Announce Type: replace Abstract: Kolmogorov--Arnold Networks (KANs) replace linear weights with spline-based functions, offering strong expressivity but posing challenges for low-precision deployment due to heterogeneous parameter distributions.

By Kazi Ahmed Asif Fuad, Lizhong Chen
efficiencybenchmarks
More like this →
About Pricing API Newsletter Sources Privacy Terms Refunds Accessibility Provider info Contact RSS

The Flow links to publishers and never republishes their articles. Summaries are machine-generated.

v1.0.0 · bb4ee0e