Requential Coding: Pushing the Limits of Model Compression with Self-Generated Training Data
arXiv:2607. 11883v1 Announce Type: new Abstract: Compression is fundamental to intelligence.
The paper investigates the limits of the maximal coding rate reduction (MCR²) framework for out‑of‑distribution (OOD) generalisation. It shows that MCR² can lead to complete prediction failure under distribution shift, even when a perfectly stable feature is available, and that adding invariance principles from IRM or REx does not resolve this issue. The authors conclude that additional assumptions or learning principles are needed to guarantee stable OOD predictions with MCR².
arXiv:2607. 11883v1 Announce Type: new Abstract: Compression is fundamental to intelligence.
Compression is fundamental to intelligence. A model that can represent its training data as a short code has discovered regularities that enable generalization.
arXiv:2505. 11702v3 Announce Type: replace Abstract: This work develops a framework for post-training augmentation invariance, in which our goal is to add invariance properties to a pretrained network without altering its behavior on the original, non-augmented input distribution.
StableVQ introduces practical guidelines to improve training stability for vector‑quantized tokenizers used in image generation models. It addresses instability caused by the entanglement of encoder–decoder and codebook training by proposing three techniques: Dynamic STE for the encoder, Region VQ Loss for the codebook, and a Decoupled Schedule for independent learning rates. Experiments on ImageNet show consistent gains in stability, codebook utilization, and reconstruction quality across various settings.
Methods operating on Vision Transformer (ViT) feature spaces typically rely on Euclidean distance or cosine similarity. This assumes that every direction is equally meaningful, but there is no reason...
arXiv:2607. 05791v1 Announce Type: cross Abstract: Boosting is a fundamental technique for generically improving the accuracy of learning algorithms (Schapire 1989).
arXiv:2606. 04857v1 Announce Type: new Abstract: Standard IMVC evaluation retrains separate models for different missing-data configurations.
arXiv:2609.27988v1 Announce Type: cross Abstract: Methods operating on Vision Transformer (ViT) feature spaces typically rely on Euclidean distance or cosine similarity. This assumes that every direc...
arXiv:2506. 14194v2 Announce Type: replace Abstract: We present a theory for the construction of out-of-distribution (OOD) detection features for neural networks.
arXiv:2606. 16196v1 Announce Type: new Abstract: Deep neural networks have achieved remarkable performance across medical imaging tasks, yet their tendency to overgeneralize under distributional shifts poses a major obstacle to safe clinical deployment.
arXiv:2606. 16050v1 Announce Type: cross Abstract: Robust deep learning under heavy-tailed and impulsive noise remains challenging because conventional losses such as mean squared error (MSE) exhibit unbounded sensitivity to outliers.
arXiv:2605. 27991v2 Announce Type: replace-cross Abstract: Gradient-flow optimization is usually viewed as an algorithmic procedure for minimizing empirical loss, with training duration selected by validation or heuristic early-stopping rules.