arXiv Machine Learning By James Li, Philip H. W. Leong, Thomas Chaffey

Quantization Robustness of Monotone Operator Equilibrium Networks

Read the original on arXiv Machine Learning →

arXiv:2603. 10562v2 Announce Type: replace-cross Abstract: Monotone operator equilibrium networks are implicit-layer models whose output is the unique equilibrium of a monotone operator, guaranteeing existence, uniqueness, and convergence.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv AI
Sep 2

Towards Provable and Scalable Training of Quantized Neural Networks with Ising Optimization

The paper presents a Quadratic Constrained Binary Optimization (QCBO) framework that provides provable guarantees for training quantized neural networks. It characterizes the topology of zero‑loss level sets, compiles finite‑depth architectures into bounded QCBOs, and introduces a sample‑wise Decomposed Lower‑Bound Optimization (DLBO) to scale Ising‑based optimization. Experiments on a coherent Ising machine show high accuracy on binary Fashion‑MNIST at 1.1‑bit precision and validate the approach on multi‑class datasets.

By Wenxin Li, Chuan Wang, Hongdong Zhu, Qi Gao, Yin Ma, Hai Wei, Kai Wen
arXiv AI
Jun 4

Model-Preserving Adaptive Rounding

arXiv:2505. 22988v3 Announce Type: replace-cross Abstract: The goal of quantization is to produce a compressed model whose output distribution is as close to the original model's as possible.

By Albert Tseng, Zhaofeng Sun, Christopher De Sa
arXiv Machine Learning
Sep 11

Why Does Post-Training Quantization Work?

Post‑training quantization compresses large language models by storing weights at reduced precision, introducing errors into hidden states that could accumulate with depth. However, pretrained models accumulate far less hidden‑state error than randomly initialized ones, largely preserving downstream performance. The study identifies two key mechanisms: (1) each layer’s new error tends to oppose inherited error, partially canceling it, and (2) the LM‑head geometry preserves high‑rank token scores, mitigating output changes.

By Yuxiang Chen, Michael Beyer, Jun Zhu, Jianfei Chen
arXiv Machine Learning
1d ago

Q-MINO: A Minimal-Norm Method for Quantization-Aware Training

The paper introduces Q-MINO, a Quantization-Aware Minimal-Norm Optimizer designed to improve training of ultra-low-bit neural networks. Q-MINO uses a temporal bundle method that incorporates gradient consensus, state-drift regularization, and an alignment constraint to produce stabilized, minimum-norm update directions. The authors solve the resulting constrained subproblem with a warm-started Frank–Wolfe procedure and provide theoretical convergence guarantees via a stochastic Lyapunov Kurdyka–Łojasiewicz framework, along with numerical experiments demonstrating its effectiveness across various quantization levels.

By Don Li