Hugging Face Trending Papers

Beyond receptive fields: sequence-pooled normalization can supply most of a sequence labeler's context

The paper argues that a convolutional sequence labeler’s receptive field is not the sole source of context; a normalization layer that computes statistics across the entire sequence at inference creates a global context path that bypasses the receptive field. By analyzing the Jacobian of the normalization layer, the authors show that this path provides almost all the benefit of a larger receptive field, especially when labels appear in long runs. Experiments on synthetic data, simulated genomes, and real 1000 Genomes haplotypes demonstrate that closing this path can multiply the value of enlarging the receptive field by up to an order of magnitude, and that ablating receptive‑field‑enlarging blocks overestimates their contribution due to this hidden path.

arXiv Machine Learning
Aug 20

Beyond receptive fields: sequence-pooled normalization can supply most of a sequence labeler's context

The paper argues that a convolutional sequence labeler’s receptive field is not the sole source of context. When a normalization layer computes statistics across the entire sequence at inference, it creates a sequence‑spanning path that supplies global context, effectively replacing the need for a larger receptive field. Experiments on synthetic labeling tasks, simulated genomes, and real 1000 Genomes haplotypes show that this global summary can match or exceed the performance of larger receptive fields, and that ablating receptive‑field‑enlarging blocks overestimates their importance.

By Qing Tian
arXiv Machine Learning
Sep 24

The Computational Value of Sensory-Aligned Receptive Fields Depends on Neuronal Expressivity

The study investigates whether sensory-aligned receptive fields provide computational benefits beyond mere resource efficiency in recurrent networks of Expressive Leaky Memory neurons. Across auditory and event-based visual classification tasks, receptive fields aligned with task-relevant sensory coordinates improve test accuracy compared to budget-matched random fields, but this advantage disappears when coordinates are scrambled or irrelevant. The benefit diminishes as neuronal expressivity increases, and generic synaptic sparsity regularization only partially recovers performance, indicating that structured receptive fields act as a computational prior beyond sparsity alone.

By Agnese Adorante, Aaron Spieler, Anna Levina
arXiv AI
Sep 1

Forget or Fine-tune? A Comparative Study of Machine Unlearning Strategies for Noisy Label Correction

The paper compares five machine unlearning (MU) methods—NegGrad, Fine‑Tuning (FT), Random Labeling (RL), SalUn, and MUNBa—on noisy‑label correction across CIFAR‑10, CIFAR‑100, and Food‑101N. Results show that the best MU strategy depends on the noise type: FT works well for most closed‑set noise, RL and SalUn are robust and nearly match retraining accuracy under instance‑dependent noise, while MUNBa excels only under extreme symmetric noise. In open‑set noise, retraining on the cleaned data actually hurts performance, indicating that approximating retraining is not suitable in that regime, yet all MU methods still achieve near‑retraining accuracy on Food‑101N with much lower runtime.

By Jo\~ao L. P. Santana, Filipe R. Cordeiro
arXiv Machine Learning
Sep 16

Attention Mean Fields Predict Average Representation Dynamics and Reveal Context-Specific Computation

The paper presents a mean‑field analysis of attention in language models, defining an average attention kernel that propagates representations layer by layer. When conditioned on a whole corpus, the kernel predicts the average evolution of representation geometry; when conditioned on a single context, it predicts the expected geometry for that context. The difference between actual attention and the mean‑field prediction—called the mean‑field deviation—captures context‑specific computation, revealing how models diverge from average behavior during training and in few‑shot tasks.

By Micah Adler, John W. Byers, Mark Crovella
arXiv AI
Sep 24

What Changed? Drift Detection with Real, Virtual, and Incomparable Diagnosis

The paper investigates drift detection in deep learning models, showing that sharing a deep encoder alone does not eliminate confounding in task-comparison scores. By introducing a conditional two‑discriminator discrepancy into the embedding space, the authors create a two‑axis gate that remains stable under input rotations and accurately tracks label‑permutation drift. This approach outperforms traditional exchange or novelty triggers, achieving high AUROC in distinguishing semantic novelty from photometric shift across multiple backbones and datasets.

By Kentaro Oda
arXiv AI
Sep 24

A Shared Encoder Is Not a Shared Task: Conditional Comparison for Deep Expert Pools

The paper demonstrates that sharing a deep encoder alone does not eliminate the confounding effects in task-comparison scores. By introducing a conditional two‑discriminator discrepancy within the embedding space, the authors achieve robust detection of task changes, maintaining stability under input rotations and accurately tracking label‑permutation drift. This approach, integrated into a mixture‑of‑heads framework, outperforms traditional novelty triggers and generalizes across multiple backbones and datasets, including ImageNet‑21k ViT‑B/16, DINOv2, and CIFAR‑100.

By Kentaro Oda