arXiv Machine Learning

Same Bit Width, Different Outcomes: Post-Training Quantization of Text-to-Speech Across Architectures

The paper evaluates post‑training quantization (PTQ) for text‑to‑speech (TTS) models across multiple architectures using a unified protocol. It shows that reducing weights to 4‑bit per‑channel can significantly lower predicted mean opinion scores (UTMOS) and that even 8‑bit per‑tensor scaling can cause severe degradation, with the impact varying by model. A staged ablation identifies the sensitive components, and per‑layer GPTQ can recover performance to within 0.1 UTMOS, while real int8 and int4 kernels confirm the simulated results on hardware, demonstrating that each configuration must be validated on the target runtime.

arXiv Computation and Language
Sep 24

Text Scores Can Miss Waveform Use: A Qwen2-Audio Quantization Case Study

The paper presents a new evaluation protocol for post‑training quantization of speech language models that separates lexical output, transcript‑insufficient endpoints, and packed implementations. In a Qwen2‑Audio case study, a 6‑bit allocation selected for translation improves chrF scores but degrades emotion recognition, while uniform and front‑layer controls perform better on emotion tasks. Similar patterns hold at 7 bits, and a 4‑bit study shows consistent emotion deficits across all low‑bit allocations, with no advantage for the selected scheme. The study highlights a precision‑dependent mismatch between lexical output, waveform‑dependent behavior, and nominal precision, without claiming a general failure of low‑bit models or a deployment benefit for the selected allocation.

By Mengzhe Geng, Jinxi Jin, Junhao Xu
arXiv Machine Learning
Sep 10

Squeeze10-LLM: Squeezing LLMs' Weights by 10 Times via a Staged Mixed-Precision Quantization Method

Squeeze10-LLM is a staged mixed‑precision post‑training quantization framework that reduces 16‑bit LLM weights to an average of 1.6 bits per weight by assigning 80% of weights to 1 bit and 20% to 4 bits. It introduces Post‑Binarization Activation Robustness (PBAR), a weight significance metric that considers activation impact, and Full Information Activation Supervision (FIAS), a strategy that preserves activation information to limit error propagation. Experiments on LLaMA and LLaMA2 demonstrate that Squeeze10‑LLM achieves state‑of‑the‑art performance for sub‑2‑bit weight‑only quantization, raising average accuracy from 43% to 56% on six zero‑shot classification tasks.

By Qingcheng Zhu, Yangyang Ren, Linlin Yang, Yanjing Li, Sheng Xu, Haodong Zhu, Juan Zhang, Runqi Wang, Baochang Zhang
arXiv AI
Sep 18

Scaling Audio Models Efficiently: Joint Optimization of Scale, Resolution, Adaptation, Precision, and Sparsity

The paper introduces a compression framework for the Whisper automatic speech recognition model that jointly optimizes six deployment dimensions—model size, temporal resolution, encoder token stride, low‑rank adaptation capacity, weight precision, and sparsity pattern—using NSGA‑III. The optimization targets three objectives: word error rate, inference FLOPs, and memory footprint. Evaluating 1,680 configurations, the study identifies compression combinations that outperform single‑axis scaling and notes that 1:4 structured sparsity cannot maintain acceptable accuracy within the tested budgets.

By Vyom Agarwal, Mokshda Gangrade, Siddharth Pal, Jerry Wu
arXiv Machine Learning
2d ago

Low-Bit Recurrent States in Hybrid Language Models

The paper introduces a method for quantizing the fixed‑size recurrent states of hybrid language models to as few as four bits per token. By deriving distortion weights from the observability Gramian and combining them with normalized state ranges, the authors achieve mixed‑precision bit allocation without requiring calibration data, rotation, or additional training. The approach also logarithmically quantizes decay rates, yielding significant reductions in excess negative log‑likelihood—up to 27.9× better than seven baselines—while maintaining near‑FP32 performance at six bits.

By Hongren Chen, Jiayang He