arXiv:2607. 02893v1 Announce Type: new Abstract: Low-bit quantization shrinks language models but treats precision as a single global hyper-parameter: every weight uses the same bit-width.
By Hamish Ogilvy
arXiv:2608.30564v1 Announce Type: cross
Abstract: Mixed-precision quantization (MPQ) assigns a different bitwidth to each linear layer of a large language model (LLM) to minimize the quantization-ind...
By Deokjae Lee, Sihun Chu, Hyun Oh Song
The paper presents a Quadratic Constrained Binary Optimization (QCBO) framework that provides provable guarantees for training quantized neural networks. It characterizes the topology of zero‑loss level sets, compiles finite‑depth architectures into bounded QCBOs, and introduces a sample‑wise Decomposed Lower‑Bound Optimization (DLBO) to scale Ising‑based optimization. Experiments on a coherent Ising machine show high accuracy on binary Fashion‑MNIST at 1.1‑bit precision and validate the approach on multi‑class datasets.
By Wenxin Li, Chuan Wang, Hongdong Zhu, Qi Gao, Yin Ma, Hai Wei, Kai Wen
arXiv:2605.11222v2 Announce Type: replace
Abstract: Quantization is an effective strategy to reduce the storage and computation footprint of large language models (LLMs). Post-training quantization (...
By Ryan Lucas, Mehdi Makni, Xiang Meng, Adam Deng, Rahul Mazumder
Cross-Backend QIEO is a runtime core for quantum-inspired evolutionary optimization that unifies execution across OpenMP5, CUDA, HIP, and multiple high-level languages. It compiles a single C++ implementation per hardware target and dispatches to CPU, multi-core, NVIDIA, or AMD backends at runtime, adapting kernels to each device’s memory hierarchy. The framework is validated with real-world bindings: a Python neural‑network hyperparameter optimizer achieving 88.60 % MNIST accuracy, a MATLAB wind‑farm layout optimizer matching particle swarm results and outperforming genetic algorithms, and a Julia package that reduces mean SSE by 2.1× in Lotka–Volterra parameter estimation.
By Aman Mittal, Ferdin Sagai Don Bosco, Kasturi Venkata Srikanth, Abhishek Singh, Aditya Singh, Abhishek Chopra
REAL-Q introduces a new post‑training quantization approach for large language models that replaces the traditional single closed‑form second‑order solver with a fine‑grained, dynamic block‑wise gradient descent applied after every 128‑column block. By aligning the surrogate loss with the end‑to‑end objective and using a sliding window for smooth cross‑layer transitions, REAL‑Q mitigates error propagation and information misalignment. Experiments on LLaMA‑3.1 and Qwen3 show up to ~49% reduction in end‑to‑end KL divergence compared to state‑of‑the‑art methods.
By Qian Zhang, Yaoming Li, Zhewen Tan, Yanshu Wang, Heng Lu, Kun Su, Zongwei Lv, Wenhan Yu, Yongge Ma, Yinjun Han, Ruikuang Liu, Tong Yang