Hugging Face Trending Papers

ExTernD: Expanded-Rank Ternary Decomposition Ternary LLM PTQ with Accuracy Approaching Any Quantization Level

We introduce ExTernD (Expanded-rank Ternary Decomposition), a post-training factorization of each LLM weight matrix $A \in \mathbb{R}^{m \times n}$ into $A \approx B \mathrm{diag}(D) C$ with ternary factors $B \in \{-1,0,+1\}^{m \times k}$, $C \in \{-1,0,+1\}^{k \times n}$ and a real scale vector $D \in \mathbb{R}^k$. The inner rank $k = μ\min(m,n)$ is deliberately expanded beyond full rank ($μ> 1$), so that components past full rank correct the quantization error of earlier ones.

arXiv AI
Jul 16

ExTernD: Expanded-Rank Ternary Decomposition Ternary LLM PTQ with Accuracy Approaching Any Quantization Level

arXiv:2607. 13511v1 Announce Type: cross Abstract: We introduce ExTernD (Expanded-rank Ternary Decomposition), a post-training factorization of each LLM weight matrix $A \in \mathbb{R}^{m \times n}$ into $A \approx B \mathrm{diag}(D) C$ with ternary factors $B \in \{-1,0,+1\}^{m \times k}$, $C \in \{-1,0,+1\}^{k \times n}$ and a real scale vector $D \in \mathbb{R}^k$.

By Chethan Reddy G. P
arXiv AI
3d ago

ShamAN-Q: Shampoo Augmented NanoQuant for Sub-1-bit LLM Weights

ShamAN-Q is a sub‑1‑bit post‑training quantization technique that builds on NanoQuant by replacing its diagonal reconstruction geometry with a dense curvature metric inspired by the Shampoo optimizer. For each linear weight, it fits a Kronecker product to the empirical Fisher information matrix of a small calibration set via Kullback–Leibler minimization, yielding a Mahalanobis reconstruction loss. The method updates continuous ADMM steps to Sylvester equations while keeping the discrete projection and deployment format unchanged, and it redistributes uniform rank across layers, achieving lower perplexity on Qwen3‑Base at roughly 1 bpw and matching or improving zero‑shot accuracy on the Eleuther LM Evaluation Harness.

By Jonathan Mei, Sang Hyub Kim, Oliver Knitter, Chi Chen, Martin Roetteler
arXiv AI
Sep 11

Scaling Post-Training Ternarisation to Qwen3-8B Capability Retention, Reproduction, Lossless Packing, and Packed Execution

The paper reports a large‑scale post‑training ternarisation of the Qwen3 language model, extending a conversion pipeline from the 4B to the 8B variant. Using KOTMS rotation, E2M‑ATQ adaptive ternarisation, and GPTQ‑style error compensation, the authors achieve a 1.361× perplexity ratio across three corpora and retain 78.5% of the FP16 accuracy on zero‑shot tasks, with the 8B model outperforming the 4B by 8.9 percentage points. The study also demonstrates lossless lattice‑aware packing, producing an 8.24 GiB checkpoint that preserves perplexity, and shows that direct packed execution can reach 15.52 tokens/s in 7.35 GiB, though packed GEMV remains slower than FP16 cuBLAS.

By Anirudh Malik, M Sparsh Mehra, Poojith Devan
arXiv Machine Learning
Sep 22

Global Ranks Survive, Selected Heads Shift: BOS-Sink Topology under 4-bit Weight-Only Quantization

The paper investigates whether sink-aware attention head selection remains valid after 4‑bit NF4 weight‑only post‑training quantization. Using Sink Topology Consistency metrics, it finds that global rank preservation stays high across Qwen2.5 and Llama‑3.2 models, yet top‑k head overlap drops to 61–79% and layer‑specific sink‑mass shifts can be substantial. The study also shows that cross‑domain calibration degrades more than within‑domain precision and that recalibration with a small number of samples can recover most of the stability, though full‑map stability may require updating more layers.

By Kuanlin Chen, Chen-Wei Kuo, Cheng-En Ou
arXiv AI
Sep 3

Post-Training Ternarization of Qwen3-4B Capability, Effective Bit Budget, Storage Compression, and Deployment

The paper reports a post‑training ternarization of the 4‑billion‑parameter Qwen model, achieving an effective 1.641‑bit representation for 81.62 % of its weights while keeping activations at 16‑bit precision. Accuracy drops from 64.5 % to 54.7 % across ten capability tests, with uneven degradation (e.g., BoolQ 84.6 % of teacher performance, ARC‑Challenge 43.8 %). After packing the ternary planes, the model size shrinks from 8.29 GiB to 3.96 GiB with negligible change in perplexity, though inference speed is not improved.

By Anirudh Malik, M Sparsh Mehra, Poojith Devan
Hugging Face Trending Papers
5d ago

From Attention Sensitivity to Layer Role: Revisiting Mixed-Precision Quantization of Transformers

The paper investigates post‑training quantization of transformer attention blocks by optimizing a joint loss over the Q, K, V projections rather than individual weight matrices. Using this joint attention‑based objective (JAB), the authors achieve significant compression on Mistral‑7B, recovering 77‑90% of the performance gap at 3 bits, but the method fails when MLP layers are included. A role‑aware offset rule that ignores sensitivity estimates outperforms JAB on GPT‑2 and full Mistral‑7B, demonstrating that the matrix a weight belongs to is more critical than sensitivity metrics.

arXiv Machine Learning
Aug 26

Low-Rank Ternary Adaptation for Fine-Tuning Transformers

The paper introduces a ternary multiplicative adaptation technique that enables fine‑tuning of ternary transformers without dequantization. By representing discrete weight updates as a low‑rank Kronecker factorization of two small ternary matrices applied element‑wise, the method preserves the ternary domain and allows direct merging of adaptation weights. Experiments on six language and vision models, including ternarized LLaMA‑3 and ViT‑B/16, show that the approach recovers most of the performance lost to quantization and outperforms existing low‑bit and ternary baselines.

By Alexandru-Dragos Manolache, Yunqiang Li, Jan van Gemert
arXiv AI
Jul 7

HiFA4: Training-Free 4-bit FlashAttention on Ascend HIF4 NPUs for LLM Inference

arXiv:2607. 04302v1 Announce Type: cross Abstract: We present HiFA4, a post-training operator-level design that executes both QK^T and PV in FlashAttention as 4-bit HIF4 Cube GEMMs for LLM inference on Ascend NPUs, while maintaining the online softmax state in FP16.

By Hui Dong, Yanzhao Li, Jie Gao, Chunlu Li, Zhiyuan Zhang, Yupeng Sun, Zhenyuan Chen, Zhiqiang Zou