arXiv Machine Learning

A Unified Framework for Quantized and Continuous Strong Lottery Tickets

arXiv:2607. 03860v1 Announce Type: new Abstract: The Strong Lottery Ticket Hypothesis (SLTH) asserts that sufficiently overparameterized, randomly initialized neural networks contain sparse subnetworks that, even without any training, can match the performance of a small trained network on a given dataset.

arXiv AI
Jun 12

Structured vs. Unstructured Pruning: An Exponential Gap

arXiv:2603. 02234v3 Announce Type: replace-cross Abstract: The Strong Lottery Ticket Hypothesis (SLTH) states that large, randomly initialized neural networks contain sparse subnetworks capable of approximating a target function at initialization without training, suggesting that pruning alone is sufficient.

By Davide Ferre' (CNRS, COATI, UniCA, I3S), Fr\'ed\'eric Giroire (I3S, COATI, UniCA), Frederik Mallmann-Trenn (CNRS, COATI, I3S, UniCA), Emanuele Natale (CNRS, COATI, I3S, UniCA)
arXiv Machine Learning
5d ago

Generalization behavior of OPTQ and the role of regularization

The paper investigates the generalization behavior of the OPTQ quantization algorithm and its stochastic variant. It derives bounds on the expected squared error when a test point is drawn from a fixed distribution, linking this error to the calibration dataset and to the regularization parameter λ. The authors use these theoretical insights to propose a new recommendation for choosing λ, which shows improved performance in experiments compared to previous suggestions.

By Erin George, Rayan Saab
Hugging Face Trending Papers
Aug 19

A Unifying Relational Perspective on Expressive Lottery Tickets

The paper investigates how sparsity impacts the expressivity of graph neural networks, focusing on relational and temporal variants. It extends the Strong Expressive Lottery Ticket Hypothesis to multi-relational and temporal domains, proving that sufficiently large RGNNs contain sparse subnetworks that preserve 1‑RWL expressivity and providing a probabilistic bound for random pruning. Experiments validate the theoretical bounds, compare them to empirical results on synthetic data, and explore the relationship between pre‑training expressivity, optimization behavior, and prediction quality on temporal and molecular benchmarks.

arXiv Machine Learning
Aug 20

A Unifying Relational Perspective on Expressive Lottery Tickets

The paper extends the Strong Expressive Lottery Ticket Hypothesis to relational and temporal graph neural networks by proving that sufficiently large RGNNs contain sparse subnetworks preserving 1‑relational Weisfeiler‑Leman expressivity. It derives a probabilistic lower bound for random pruning to achieve such subnetworks and shows that common TGNNs and cross‑graph message passing can be reformulated as RGNNs to inherit these guarantees. Experiments validate the bound, compare it to empirical probabilities on synthetic data, and explore the relationship between pre‑training expressivity, optimization behavior, and prediction quality on temporal and molecular benchmarks.

By Lorenz Kummer, Samir Moustafa, Anatol Ehrlich, Franka Bause, Marco Nennstiel, Przemys{\l}aw Andrzej Wa{\l}\c{e}ga, Nils Morten Kriege
arXiv AI
Jun 4

Model-Preserving Adaptive Rounding

arXiv:2505. 22988v3 Announce Type: replace-cross Abstract: The goal of quantization is to produce a compressed model whose output distribution is as close to the original model's as possible.

By Albert Tseng, Zhaofeng Sun, Christopher De Sa
arXiv AI
Sep 2

Towards Provable and Scalable Training of Quantized Neural Networks with Ising Optimization

The paper presents a Quadratic Constrained Binary Optimization (QCBO) framework that provides provable guarantees for training quantized neural networks. It characterizes the topology of zero‑loss level sets, compiles finite‑depth architectures into bounded QCBOs, and introduces a sample‑wise Decomposed Lower‑Bound Optimization (DLBO) to scale Ising‑based optimization. Experiments on a coherent Ising machine show high accuracy on binary Fashion‑MNIST at 1.1‑bit precision and validate the approach on multi‑class datasets.

By Wenxin Li, Chuan Wang, Hongdong Zhu, Qi Gao, Yin Ma, Hai Wei, Kai Wen