arXiv:2606. 00289v1 Announce Type: new Abstract: Quantization is a fundamental tool used to compress datasets, neural network weights, and memory usage in a range of computational tasks.
By Nathan White, Krish Singal
arXiv:2510. 18784v3 Announce Type: replace Abstract: Despite significant work on low-bit quantization-aware training (QAT), there is still an accuracy gap between such techniques and native training.
By Soroush Tabesh, Mher Safaryan, Andrei Panferov, Alexandra Volkova, Dan Alistarh
arXiv:2609.06430v1 Announce Type: new
Abstract: We study the identity straight-through estimator (STE) for training a two-layer binary-activation network with hinge loss from the perspective of Stati...
By Yiming Ying
arXiv:2608. 01357v1 Announce Type: new Abstract: Traditional approximation theory measures convergence rates in terms of the number of parameters or degrees of freedom.
By Tong Mao, Jinchao Xu
arXiv:2510. 10101v4 Announce Type: replace Abstract: Understanding the interplay between generalization, expressivity, and the geometry of the input space is a central challenge in graph learning.
By Martin Carrasco, Caio F. Deberaldini Netto, Vahan A. Martirosyan, Ehimare Okoyomon, Caterina Graziani
arXiv:2505. 22988v3 Announce Type: replace-cross Abstract: The goal of quantization is to produce a compressed model whose output distribution is as close to the original model's as possible.
By Albert Tseng, Zhaofeng Sun, Christopher De Sa
The paper proves that training a binary quantized neural network (2-QNNT) is W[1]-hard when parameterized solely by the sum of input and output dimensions, α+ω. This hardness result holds even for zero training error on a specially constructed dataset where each input equals its target and the examples form a coordinate‑wise prefix chain. The proof reduces from DAG edge‑disjoint paths, employing a one‑flip routing equivalence that links activation transitions to vertex‑disjoint paths in the network.
By Tao Jiang, Minbo Gao, Shaowei Cai
arXiv:2608. 18147v1 Announce Type: cross Abstract: Adaptive stochastic quantization (ASQ) is a recently introduced quantization approach that optimizes the Mean Squared Error (MSE) for a given input while preserving unbiasedness.
By Ran Ben Basat, Yaniv Ben-Itzhak, Michael Mitzenmacher, Shay Vargaftik
The paper investigates the generalization behavior of the OPTQ quantization algorithm and its stochastic variant. It derives bounds on the expected squared error when a test point is drawn from a fixed distribution, linking this error to the calibration dataset and to the regularization parameter λ. The authors use these theoretical insights to propose a new recommendation for choosing λ, which shows improved performance in experiments compared to previous suggestions.
By Erin George, Rayan Saab
arXiv:2206. 04359v3 Announce Type: replace Abstract: One of the fundamental challenges in the deep learning community is to theoretically understand how well a deep neural network generalizes to unseen data.
By Chengli Tan, Jiangshe Zhang, Junmin Liu, Yihong Gong
arXiv:2503. 07325v2 Announce Type: replace Abstract: Understanding and certifying the behavior of modern deep neural networks remains a fundamental challenge in reliable machine learning.
By Khoat Than, Dat Phan
arXiv:2607. 03860v1 Announce Type: new Abstract: The Strong Lottery Ticket Hypothesis (SLTH) asserts that sufficiently overparameterized, randomly initialized neural networks contain sparse subnetworks that, even without any training, can match the performance of a small trained network on a given dataset.
By Aakash Kumar, Emanuele Natale