arXiv Machine Learning By Osama Orabi, Artur Zagitov, Hadi Salloum, Viktor A. Lobachev, Yaroslav Kholodov

Task-Aware QUBO Allocation for Mixed-Precision Quantization

Read the original on arXiv Machine Learning →

The paper introduces a task‑aware quadratic unconstrained binary optimization (QUBO) surrogate for mixed‑precision quantization, separating weight and activation profiles and incorporating a bit‑operation (BOP) cost and structural priors. Using this surrogate, a network‑wide allocation is refined via a direct validation‑based PROTES search, achieving a 37.192 dB PSNR on a compact NAFBlock denoiser with 4.035% routed‑layer BOPs, slightly better than a HAWQ‑style baseline. The study also shows that LSQ+ refinement narrows the quality gap and that the benefit of refinement varies with architecture and recovery strategy.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv Machine Learning
Jul 17

PolyQ: Codesigning End-to-End Quantization Framework for Scalable Edge CPU LLM Inference

arXiv:2607. 14618v1 Announce Type: new Abstract: CPUs are the most universal target for on-device LLM inference, but existing low-bit quantization methods offer either coarse operating points or fine-grained mixed precision that is difficult to execute efficiently on CPUs.

By Hyunwoo Oh, Suyeon Jang, Hanning Chen, KyungIn Nam, Sanggeon Yun, Ryozo Masukawa, Mohsen Imani
arXiv Machine Learning
Sep 14

Dynamic Expert Quantization for Scalable Mixture-of-Experts Inference

Dynamic Expert Quantization (DynaExq) is a runtime-aware mixed-precision serving system designed for single‑GPU Mixture‑of‑Experts (MoE) inference under a hard high‑bandwidth memory (HBM) envelope. It treats the problem as an online, budget‑constrained precision allocation task, keeping the most frequently used experts at higher precision while relegating the rest to low‑precision fallbacks. By estimating expert hotness from router traces and asynchronously promoting or demoting experts, DynaExq maintains a fully materialized expert set during the forward pass, improving accuracy and throughput compared to static post‑training quantization and offloading/prefetch baselines. whyItMatters":"DynaExq enables efficient deployment of large MoE models on memory‑limited GPUs by dynamically allocating precision based on runtime expert usage, thereby reducing memory footprint and latency while boosting accuracy and throughput."

By Kexin Chu, Dawei Xiang, Zixu Shen, Yiwei Yang, Zecheng Liu, Wei Zhang