arXiv Machine Learning By Cristhian Kapelinski, Diego Kreutz

Not All 4-bit Quantizers Are Equal: Deployment-Time Mitigation of PII Leakage in Fine-Tuned Small Language Models

Read the original on arXiv Machine Learning →

The paper investigates how different 4‑bit quantization techniques affect privacy when fine‑tuned small language models are deployed. It finds that methods using a calibration corpus, such as Activation‑aware Weight Quantization (AWQ) and Gradient‑based Post‑Training Quantization (GPTQ), prevent the reproduction of planted private records, whereas a calibration‑free format (GGUF Q4_K_M) leaks 5.3% of them. Across models ranging from 0.5 to 7 billion parameters, AWQ consistently leaks the least while maintaining minimal accuracy loss, indicating that the choice of 4‑bit method is a privacy decision as well as a performance one.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv AI
Aug 10

An Empirical Study of openPangu Quantization on Ascend NPUs

arXiv:2606. 21257v4 Announce Type: replace-cross Abstract: openPangu models are attractive targets for private and domestic large-language-model deployment, yet their robustness under aggressive post-training quantization on Ascend NPUs has not been systematically characterized.

By Tong Shi, Jiacheng Wang, Hui Xie, Ying Li, Aishan Liu, Jinyang Guo, Xianglong Liu
arXiv AI
Jun 29

An Empirical Study of OpenPangu Quantization on Ascend NPUs

arXiv:2606. 21257v2 Announce Type: replace-cross Abstract: OpenPangu models are attractive targets for private and domestic large-language-model deployment, yet their robustness under aggressive post-training quantization on Ascend NPUs has not been systematically characterized.

By Tong Shi, Jiacheng Wang, Hui Xie, Ying Li, Aishan Liu, Jinyang Guo, Xianglong Liu
arXiv Machine Learning
Sep 10

Squeeze10-LLM: Squeezing LLMs' Weights by 10 Times via a Staged Mixed-Precision Quantization Method

Squeeze10-LLM is a staged mixed‑precision post‑training quantization framework that reduces 16‑bit LLM weights to an average of 1.6 bits per weight by assigning 80% of weights to 1 bit and 20% to 4 bits. It introduces Post‑Binarization Activation Robustness (PBAR), a weight significance metric that considers activation impact, and Full Information Activation Supervision (FIAS), a strategy that preserves activation information to limit error propagation. Experiments on LLaMA and LLaMA2 demonstrate that Squeeze10‑LLM achieves state‑of‑the‑art performance for sub‑2‑bit weight‑only quantization, raising average accuracy from 43% to 56% on six zero‑shot classification tasks.

By Qingcheng Zhu, Yangyang Ren, Linlin Yang, Yanjing Li, Sheng Xu, Haodong Zhu, Juan Zhang, Runqi Wang, Baochang Zhang