arXiv Machine Learning

ESTS at WMT26: Routing-Informed Expert Pruning for Model Compression

The paper reports six submissions by the ESTS team to the WMT26 Model Compression Shared Task for English–Simplified Chinese and English–Egyptian Arabic. Each submission offers three compression operating points derived from GPT‑OSS‑20B, using routing‑informed expert pruning, cross‑lingual routing divergence for capacity allocation, and MXFP4 quantization of retained expert projection weights. The resulting models, ranging from 4.186 B to 7.770 B parameters, are fine‑tuned on GPT‑5.1 synthetic data and evaluated internally with xCOMET‑XL.

arXiv Machine Learning
Aug 11

Statistically-Lossless Quantization of Large Language Models

arXiv:2605. 02404v2 Announce Type: replace Abstract: Model quantization has become essential for efficient large language model deployment, yet existing approaches present clear trade-offs: methods such as GPTQ and AWQ achieve practical compression but are lossy, while lossless techniques preserve fidelity but lack inference acceleration.

By Michael Helcig, Eldar Kurtic, Dan Alistarh
arXiv Machine Learning
Aug 28

LowRankArena: A Standardized Evaluation Platform for SVD-Based LLM Compression

LowRankArena is a standardized evaluation platform for SVD‑based low‑rank compression of large language models, unifying task versions, compression budgets, comparison regimes, and inference measurements. It provides a reproducible pipeline with over 3 TiB of released compressed checkpoints, enabling consistent comparisons across methods. An audit of five representative SVD techniques using LowRankArena shows that prior reported gains are highly conditional, with performance leaders and tiers shifting across backbones and keep ratios, and that nominal low‑rank savings often yield limited end‑to‑end speedups.

By Zishan Shao, Lixun Zhang, Kangning Cui, Wenhao Wu, Jinhee Kim, Yixiao Wang, Ting Jiang, Hancheng Ye, Qinsi Wang, Fan Yang, Danyang Zhuo, Yiran Chen, Hai Li