arXiv Machine Learning

Q-DEQ: Discrete Solving and Quantization for Deep Equilibrium Models in Time Series Forecasting under Edge Deployment Coding Constraints

arXiv AI
Sep 2

Towards Provable and Scalable Training of Quantized Neural Networks with Ising Optimization

The paper presents a Quadratic Constrained Binary Optimization (QCBO) framework that provides provable guarantees for training quantized neural networks. It characterizes the topology of zero‑loss level sets, compiles finite‑depth architectures into bounded QCBOs, and introduces a sample‑wise Decomposed Lower‑Bound Optimization (DLBO) to scale Ising‑based optimization. Experiments on a coherent Ising machine show high accuracy on binary Fashion‑MNIST at 1.1‑bit precision and validate the approach on multi‑class datasets.

By Wenxin Li, Chuan Wang, Hongdong Zhu, Qi Gao, Yin Ma, Hai Wei, Kai Wen
arXiv Computation and Language
6d ago

Cross-Backend QIEO: Universal Runtime Portability across OpenMP5, CUDA, HIP, and Multi-Language Interfaces

Cross-Backend QIEO is a runtime core for quantum-inspired evolutionary optimization that unifies execution across OpenMP5, CUDA, HIP, and multiple high-level languages. It compiles a single C++ implementation per hardware target and dispatches to CPU, multi-core, NVIDIA, or AMD backends at runtime, adapting kernels to each device’s memory hierarchy. The framework is validated with real-world bindings: a Python neural‑network hyperparameter optimizer achieving 88.60 % MNIST accuracy, a MATLAB wind‑farm layout optimizer matching particle swarm results and outperforming genetic algorithms, and a Julia package that reduces mean SSE by 2.1× in Lotka–Volterra parameter estimation.

By Aman Mittal, Ferdin Sagai Don Bosco, Kasturi Venkata Srikanth, Abhishek Singh, Aditya Singh, Abhishek Chopra
arXiv AI
Sep 2

REAL-Q: E2E LLM Quantization via Dynamic Gradient Descent

REAL-Q introduces a new post‑training quantization approach for large language models that replaces the traditional single closed‑form second‑order solver with a fine‑grained, dynamic block‑wise gradient descent applied after every 128‑column block. By aligning the surrogate loss with the end‑to‑end objective and using a sliding window for smooth cross‑layer transitions, REAL‑Q mitigates error propagation and information misalignment. Experiments on LLaMA‑3.1 and Qwen3 show up to ~49% reduction in end‑to‑end KL divergence compared to state‑of‑the‑art methods.

By Qian Zhang, Yaoming Li, Zhewen Tan, Yanshu Wang, Heng Lu, Kun Su, Zongwei Lv, Wenhan Yu, Yongge Ma, Yinjun Han, Ruikuang Liu, Tong Yang
arXiv Machine Learning
Sep 24

Task-Aware QUBO Allocation for Mixed-Precision Quantization

The paper introduces a task‑aware quadratic unconstrained binary optimization (QUBO) surrogate for mixed‑precision quantization, separating weight and activation profiles and incorporating a bit‑operation (BOP) cost and structural priors. Using this surrogate, a network‑wide allocation is refined via a direct validation‑based PROTES search, achieving a 37.192 dB PSNR on a compact NAFBlock denoiser with 4.035% routed‑layer BOPs, slightly better than a HAWQ‑style baseline. The study also shows that LSQ+ refinement narrows the quality gap and that the benefit of refinement varies with architecture and recovery strategy.

By Osama Orabi, Artur Zagitov, Hadi Salloum, Viktor A. Lobachev, Yaroslav Kholodov
arXiv Machine Learning
Sep 24

Predicting Quantization Price for Selecting PTQ Configurations Before Deployment

The paper proposes a method for selecting post‑training quantization (PTQ) configurations before deployment by treating each admissible layer configuration as an error generator with an associated deployment cost. It introduces a priced layer‑output error framework that uses the covariance of layer outputs and the full‑precision model’s curvature to compute a price for each configuration. This approach replaces traditional reconstruction or Hessian‑based scores with a unified, cost‑aware selector that can calibrate and budget PTQ settings efficiently.

By Junbin Qiu, Jian Mu, Weitong Zhang, Yao Shu
arXiv Machine Learning
Aug 28

A Layer Importance Metric for Quantization Accounting for the Speed-Quality Trade-off in Autoregressive Models

The paper introduces a composite metric for quantizing small language models that balances information retention and throughput gains, using a normalized SQNR-based coefficient and roofline-based latency analysis. Profiling Gemma 3 1B shows that Feed‑Forward Network blocks and the embedding matrix are prime candidates for acceleration, with the metric enabling tuning of speed‑quality trade‑offs without actual execution. The authors demonstrate that their estimates predict accelerated speedup within about 4% error and allocate resources more effectively than evolutionary search or Shapley‑value methods.

By Artem Safronov
arXiv Machine Learning
Sep 11

Optimizing AI Inference Across the Deployment Stack

The paper argues that AI deployment performance depends on interactions among compression, compiler transformations, and serving policies rather than just model architecture. It introduces a three‑layer taxonomy—model‑level techniques, compiler transformations, and system policies—and frames deployment as a constrained multi‑objective optimization problem over accuracy, latency, throughput, memory footprint, and energy. The authors propose an evidence protocol for comparable benchmarking and synthesize data from edge and data‑center platforms to show that cross‑layer interactions drive deployment outcomes, concluding with a constraint‑aware selection procedure and open research problems.

By Tejinder Singh, John Pflueger, Jeebak Mitra, Robert Lincourt, Mitchell Markow, Bhavesh A. Patel
arXiv Machine Learning
Jul 31

Complementary Matrix-Gated QKAN Fast-Weight Programmers for Quantum Dynamics Forecasting

arXiv:2607. 27945v1 Announce Type: cross Abstract: Sequence models must decide what to write into memory and what to retain.

By Kuo-Chung Peng, Samuel Yen-Chi Chen, Jiun-Cheng Jiang, Chen-Yu Liu, En-Jui Kuo, Yun-Yuan Wang, Tzung-Chi Huang, Prayag Tiwari, Chi-Sheng Chen, Chun-Hua Lin, Yu-Chao Hsu, Tai-Yue Li, Saif Al-Kuwari, Simon See, Kuan-Cheng Chen, Nan-Yow Chen, Hsi-Sheng Goan