Hugging Face Blog

Ettin Suite: SoTA Paired Encoders and Decoders

arXiv AI
Sep 2

Superposed Latent Autoencoder

The paper introduces the Superposed Latent Autoencoder (SLAE), a method that stores multiple wide latent representations together by superposing them into a single memory tensor using learned codes and randomized keys. SLAE eliminates the need for tight dimensional bottlenecks, achieving up to 56% lower reconstruction error on datasets such as CIFAR-10/100 and SVHN while maintaining the same storage budget. The approach also boosts downstream classification performance by up to 16.79 percentage points, demonstrating that wide representations can be effectively compressed through structured interference rather than dimensional reduction.

By Quanling Zhao, Jiaying Yang, Tianqi Zhang, Ziyang Hao, Fatemeh Asgarinejad, Flavio Ponzina, Tajana Rosing
arXiv Computer Vision
Sep 11

OmniKVQuant: KV Cache Quantization for Omni-LLMs

OmniKVQuant introduces a training‑free framework for quantizing the key‑value (KV) cache of omni‑modal large language models (Omni‑LLMs) that process audio, video, and text simultaneously. The method addresses two identified problems—temporal key drift and heterogeneous value geometry—by setting key quantization ranges over short input windows and rotating values separately for each modality. Applied to Qwen2.5‑Omni and Qwen3‑Omni, OmniKVQuant achieves 2‑bit KV caches while largely preserving performance across seven audio‑visual benchmarks, and includes a fused Triton decode kernel that eliminates the need for a dense FP16 cache.

By Suho Yoo, Hyunjong Ok, Jongmin Choi, Jihoo Jung, Joon Son Chung
arXiv Machine Learning
Sep 24

Six Layers Less: Encoder Pruning for Whisper with Label-Free Recovery

The paper introduces a method to prune six layers from the encoder of OpenAI’s Whisper ASR model, reducing the encoder stack by 18.5% without requiring custom inference code. Layers are selected based on their minimal impact on Word Error Rate when removed. After pruning, the model’s WER rises from 18.2% to 21.9%, but distillation with unlabeled monolingual speech data lowers it to 20.1%. "whyItMatters":"The approach offers a straightforward way to accelerate Whisper inference by simplifying the encoder while maintaining acceptable accuracy, and the released code and model enable immediate adoption by the community."

By Rasmus Aagaard, Nicki Skafte Detlefsen