arXiv Machine Learning

Where Quality Breaks in Compressed Short-Text Generation: Staged Bottleneck Localization

arXiv:2607. 24176v1 Announce Type: cross Abstract: Compressed short-text generators can fail in two different places: the codec may discard information before generation starts, or the latent generator may produce weak codes.

arXiv Computer Vision
3d ago

Rethinking Generative Image Compression at Extremely Low Bitrates

The paper introduces RAE-CoD, a diffusion-based compression method that operates in a representation autoencoder space to preserve recognizable content even at extremely low bitrates. It addresses the problem of semantic collapse observed in existing codecs when the bitrate approaches zero, showing that reconstruction losses conflict with semantic objectives and that VAE diffusion models lose efficiency in preserving semantics. Experiments on MSCOCO-30K demonstrate that RAE-CoD outperforms competitors, reducing VFM feature MSE and Fréchet Distance ratios by at least 25.7% and 69.1% at 0.001–0.008 bpp while maintaining stable recognizability and quality.

By Tianyu Zhang, Zhaoyang Jia, Houqiang Li, Dong Liu
arXiv Machine Learning
Aug 28

LowRankArena: A Standardized Evaluation Platform for SVD-Based LLM Compression

LowRankArena is a standardized evaluation platform for SVD‑based low‑rank compression of large language models, unifying task versions, compression budgets, comparison regimes, and inference measurements. It provides a reproducible pipeline with over 3 TiB of released compressed checkpoints, enabling consistent comparisons across methods. An audit of five representative SVD techniques using LowRankArena shows that prior reported gains are highly conditional, with performance leaders and tiers shifting across backbones and keep ratios, and that nominal low‑rank savings often yield limited end‑to‑end speedups.

By Zishan Shao, Lixun Zhang, Kangning Cui, Wenhao Wu, Jinhee Kim, Yixiao Wang, Ting Jiang, Hancheng Ye, Qinsi Wang, Fan Yang, Danyang Zhuo, Yiran Chen, Hai Li
arXiv Machine Learning
Jun 24

ParallelBench: Understanding the Trade-offs of Parallel Decoding in Diffusion LLMs

arXiv:2510. 04767v2 Announce Type: replace Abstract: While most autoregressive LLMs are constrained to one-by-one decoding, diffusion LLMs (dLLMs) have attracted growing interest for their potential to dramatically accelerate inference through parallel decoding.

By Wonjun Kang, Kevin Galim, Seunghyuk Oh, Minjae Lee, Yuchen Zeng, Shuibai Zhang, Coleman Hooper, Yuezhou Hu, Hyung Il Koo, Nam Ik Cho, Kangwook Lee
arXiv Machine Learning
Sep 14

ESTS at WMT26: Routing-Informed Expert Pruning for Model Compression

The paper reports six submissions by the ESTS team to the WMT26 Model Compression Shared Task for English–Simplified Chinese and English–Egyptian Arabic. Each submission offers three compression operating points derived from GPT‑OSS‑20B, using routing‑informed expert pruning, cross‑lingual routing divergence for capacity allocation, and MXFP4 quantization of retained expert projection weights. The resulting models, ranging from 4.186 B to 7.770 B parameters, are fine‑tuned on GPT‑5.1 synthetic data and evaluated internally with xCOMET‑XL.

By Liu O. Martin, Lucas Bandarkar, Nanyun Peng