Hugging Face Trending Papers

Requential Coding: Pushing the Limits of Model Compression with Self-Generated Training Data

Compression is fundamental to intelligence. A model that can represent its training data as a short code has discovered regularities that enable generalization.

arXiv Machine Learning
Sep 21

On the Limits of Maximal Coding Rate Reduction for Out-of-Distribution Generalisation

The paper investigates the limits of the maximal coding rate reduction (MCR²) framework for out‑of‑distribution (OOD) generalisation. It shows that MCR² can lead to complete prediction failure under distribution shift, even when a perfectly stable feature is available, and that adding invariance principles from IRM or REx does not resolve this issue. The authors conclude that additional assumptions or learning principles are needed to guarantee stable OOD predictions with MCR².

By Menghui Zhou, Gaoshan Bi, Vitaveska Lanfranchi, Po Yang
arXiv AI
6d ago

Teacher-Anchored Selection of Post-Training Quantized Models under Domain Shift

The paper investigates how to choose the best quantized model from a family of compressed versions when target labels are scarce or unavailable. It finds that a simple rule based on minimum teacher distortion consistently selects the same eight‑bit, per‑channel, unclipped configuration, though this does not minimize empirical target cross‑entropy. The study also shows that confidence‑based estimators perform poorly in overconfident regimes, while output‑distribution estimators can outperform the teacher in some architectures, and that combining distortion with a supervised term can improve selection. Across 134 candidate families, teacher‑anchored selection reduces mean regret with very few labels, though the benefit diminishes after about 25 labels.

By Alejandro Rodriguez Dominguez, Muhammad Shahzad, Xia Hong
Hugging Face Trending Papers
Jul 28

Bridging Compute- and Data-Optimal Pretraining

Classical compute-optimal scaling laws assume an unbounded supply of fresh pretraining data, yet pretraining is increasingly entering a regime in which compute grows faster than the availability of high-quality data. We propose Compute-Data (CD) scaling laws, a unified framework that bridges compute-optimal scaling, where data scales freely with compute, and data-optimal scaling, where the corpus is fixed while compute can grow without bound.