arXiv AI By Huicheng Zhang, Xiyao Feng, Ze-Tong Li, Chengkai Zhu, Xiao Shi, Xiwei Pan, Jinguo Liu, Ge Bai, Xin Wang

Per-Matrix Optimality Is Not Enough: Three-Level Optimization for Low-Rank LLM Compression

Read the original on arXiv AI →

The Flow has not summarised this story yet — read it at arXiv AI.

arXiv Computation and Language
1d ago

Learning Functional Subspaces for Neural Network Compression

arXiv:2609.40127v1 Announce Type: cross Abstract: Modern transformers pair impressive capabilities with substantial memory and compute demands. Low-rank weight factorization reduces both while keepin...

By Massimo Bini, Anders Christensen, Stephan Alaniz, Judah Goldfeder, Ole Winther, Yann LeCun, Ravid Shwartz-Ziv, Zeynep Akata
arXiv AI
Sep 11

Scaling Post-Training Ternarisation to Qwen3-8B Capability Retention, Reproduction, Lossless Packing, and Packed Execution

The paper reports a large‑scale post‑training ternarisation of the Qwen3 language model, extending a conversion pipeline from the 4B to the 8B variant. Using KOTMS rotation, E2M‑ATQ adaptive ternarisation, and GPTQ‑style error compensation, the authors achieve a 1.361× perplexity ratio across three corpora and retain 78.5% of the FP16 accuracy on zero‑shot tasks, with the 8B model outperforming the 4B by 8.9 percentage points. The study also demonstrates lossless lattice‑aware packing, producing an 8.24 GiB checkpoint that preserves perplexity, and shows that direct packed execution can reach 15.52 tokens/s in 7.35 GiB, though packed GEMV remains slower than FP16 cuBLAS.

By Anirudh Malik, M Sparsh Mehra, Poojith Devan
arXiv Machine Learning
Sep 3

XMerge: Cross-Axis Selection and Reconstructive Layer Merging for LLM Depth Compression

XMerge is a post‑training method for compressing large language models by removing entire transformer layers while preserving a standard serving architecture. It selects low‑impact blocks via cross‑axis selection and refits adjacent surviving blocks with local boundary reconstruction, requiring no task labels or fine‑tuning. Across seven Llama and Qwen backbones, XMerge outperforms five published baselines, achieving top rankings on CORE and MMLU tasks even at aggressive compression levels, and consistently avoids model collapse while improving calibration and decoding efficiency.

By Jundong Hu, Shekar Ramachandran