arXiv Machine Learning

Low-Rank Ternary Adaptation for Fine-Tuning Transformers

The paper introduces a ternary multiplicative adaptation technique that enables fine‑tuning of ternary transformers without dequantization. By representing discrete weight updates as a low‑rank Kronecker factorization of two small ternary matrices applied element‑wise, the method preserves the ternary domain and allows direct merging of adaptation weights. Experiments on six language and vision models, including ternarized LLaMA‑3 and ViT‑B/16, show that the approach recovers most of the performance lost to quantization and outperforms existing low‑bit and ternary baselines.

Hugging Face Trending Papers
Jun 25

CAT-Q: Cost-efficient and Accurate Ternary Quantization for LLMs

In this paper, we present CAT-Q, Cost-efficient and Accurate Ternary Quantization, for compressing and accelerating LLMs. Unlike existing state-of-the-art ternary quantization methods that rely on data-intensive and costly quantization-aware training to mitigate severe performance degradation, CAT-Q is a simple yet effective post-training quantization scheme that is readily applicable to LLMs with diverse architectures and model sizes.

arXiv AI
Jul 16

ExTernD: Expanded-Rank Ternary Decomposition Ternary LLM PTQ with Accuracy Approaching Any Quantization Level

arXiv:2607. 13511v1 Announce Type: cross Abstract: We introduce ExTernD (Expanded-rank Ternary Decomposition), a post-training factorization of each LLM weight matrix $A \in \mathbb{R}^{m \times n}$ into $A \approx B \mathrm{diag}(D) C$ with ternary factors $B \in \{-1,0,+1\}^{m \times k}$, $C \in \{-1,0,+1\}^{k \times n}$ and a real scale vector $D \in \mathbb{R}^k$.

By Chethan Reddy G. P
Hugging Face Trending Papers
Jul 15

ExTernD: Expanded-Rank Ternary Decomposition Ternary LLM PTQ with Accuracy Approaching Any Quantization Level

We introduce ExTernD (Expanded-rank Ternary Decomposition), a post-training factorization of each LLM weight matrix $A \in \mathbb{R}^{m \times n}$ into $A \approx B \mathrm{diag}(D) C$ with ternary factors $B \in \{-1,0,+1\}^{m \times k}$, $C \in \{-1,0,+1\}^{k \times n}$ and a real scale vector $D \in \mathbb{R}^k$. The inner rank $k = μ\min(m,n)$ is deliberately expanded beyond full rank ($μ> 1$), so that components past full rank correct the quantization error of earlier ones.

arXiv AI
2d ago

TopK-Guided: Adaptive, Budget-Aware Activation Sparsity for Efficient LLM Inference

TopK-Guided is a training‑free method that improves activation sparsity for large language model inference by combining token‑level sparsity adaptation with block‑level budget allocation that accounts for block sensitivity. It addresses limitations of existing methods like TEAL, which adapts sparsity per token but lacks tight control, and WINA, which enforces a fixed sparsity across all tokens and blocks. Experiments on Llama‑2 and Llama‑3 show that TopK‑Guided consistently yields better perplexity and downstream accuracy while maintaining similar compute costs to WINA, especially at high sparsity levels.

By Mukund Agarwalla, Chih-Jen Lin