arXiv AI

Useful to Whom? Sample Value Is Defined Only Relative to the Learner

The paper investigates how the usefulness of training samples, as determined by coreset selection, depends on the learner rather than just the data. Experiments on ImageNet-100 and ImageNet-1k show that changing model width, input grid, stride, and architecture (e.g., ResNet vs. ViT) shifts the crossover point where different selection criteria (easy-first vs. geometric coverage) become optimal. These findings demonstrate that the relative value of a fixed subset of samples varies with the target learner’s capacity and structure, and that selection strategies must be tuned to the specific model they will train.

arXiv Machine Learning
Sep 24

The Drift Contract: Spectral Updates for Depth-Robust Local Learning

The paper introduces the Drift Contract, a spectral update geometry for local learning that improves depth robustness and hyperparameter stability. By applying momentum orthogonalization with spectral step scaling to per‑layer updates, the authors achieve consistent performance across a wide range of widths and depths on CIFAR‑10 MLPs, outperforming local Adam and providing a per‑layer, input‑conditioned drift bound. The study also shows that the spectral geometry itself, rather than step‑size rules, drives the observed depth robustness, while a negative result indicates that the stability benefit is limited to non‑normalized layers.

By Fabien Polly
arXiv Machine Learning
Sep 22

Are Coreset Selection Methods Worth Their Cost?

The paper evaluates coreset selection methods by incorporating both selection and training time into a unified wall‑clock budget, using a standardized benchmark across four datasets and multiple selectors. Across numerous budget anchors, simple random or full‑data training consistently outperforms sophisticated selectors, and selection costs are dominated by a full‑dataset scan that cannot be amortized. The study also identifies when subset reuse can justify selection and reports several correctness fixes in a popular codebase.

By Yangze Liu, Zhongyi Han
arXiv AI
Jul 22

Soft-TransFormers for Continual Learning

arXiv:2411. 16073v4 Announce Type: replace-cross Abstract: Inspired by the Well-initialized Lottery Ticket Hypothesis (WLTH), we introduce Soft-TransFormers (Soft-TF), a continual learning framework that adapts a frozen pre-trained Transformer through task-specific soft subnetworks: real-valued multiplicative masks over the query, key, value, and output projections of selected self-attention layers.

By Haeyong Kang, Chang D. Yoo
arXiv Computer Vision
Aug 24

When does fusing hand-crafted knowledge with learned representations pay? A cost-normalized benchmark of stacking, substitution, and interference

arXiv:2608.21098v1 Announce Type: new Abstract: Fusing prior knowledge with data-driven learning is attractive where data is scarce, yet no controlled account says when it helps, is redundant, or har...

By Ahmad AlMughrabi, Albert Clop, Benjamin Busam, Ricardo Marques, Petia Radeva
arXiv Computer Vision
6d ago

Training-Free Bottleneck Width Planning for Convolutional Autoencoders

The paper introduces Multiscale Spectral Rate‑Distortion (MS‑SRD), a training‑free method that predicts the required bottleneck channel width for convolutional autoencoders at user‑specified spatial cuts, using only training images and a normalized mean‑squared error bound. MS‑SRD’s covariance‑tail rule is exact for shared linear block‑convolutional autoencoders under squared error, and a nested‑scale dominance result allows reporting an activation‑parameter Pareto frontier alongside the minimal‑latent candidate. Across thirteen grayscale datasets, the method achieves a 0.84% mean absolute percentage error in latent‑size prediction, with most predictions exact or within one channel, and demonstrates comparable performance to retrospective external widths in deployable comparisons without any training of a selector.

By Guannan Guo
Hugging Face Trending Papers
Jun 24

Pre-Warm: Input-Conditioned Weight Initialization for Convolutional Neural Networks

We introduce Pre-Warm, a simple yet effective zero-training-cost method for data-conditioned initialization of the first convolutional layer. Before the first forward pass, Pre-Warm extracts mean-centered local patches from a single training batch, clusters them with MiniBatchKMeans, applies inverse Manhattan spatial weighting, and uses the resulting centroids to initialize half of the first-layer filters (the remainder retain Kaiming initialization).