arXiv Machine Learning

Coarse-Graining Hidden Representations: Unsupervised Neuron Selection via Mapping Entropy

The paper introduces an unsupervised method for selecting essential neurons in overparameterized neural networks by minimizing mapping entropy (ME), a metric that quantifies the loss of discriminatory power when neurons are discarded. ME-based selection relies solely on hidden-activation statistics and, in experiments, identifies minimal teacher-consistent representations in teacher‑student networks and coherent functional-class mappings in a non‑linear Gaussian process task. Subnetworks chosen by ME outperform random subsets of the same size, especially under strong compression, on both a Gaussian process task and translation‑augmented MNIST.

arXiv AI
6d ago

Teacher-Anchored Selection of Post-Training Quantized Models under Domain Shift

The paper investigates how to choose the best quantized model from a family of compressed versions when target labels are scarce or unavailable. It finds that a simple rule based on minimum teacher distortion consistently selects the same eight‑bit, per‑channel, unclipped configuration, though this does not minimize empirical target cross‑entropy. The study also shows that confidence‑based estimators perform poorly in overconfident regimes, while output‑distribution estimators can outperform the teacher in some architectures, and that combining distortion with a supervised term can improve selection. Across 134 candidate families, teacher‑anchored selection reduces mean regret with very few labels, though the benefit diminishes after about 25 labels.

By Alejandro Rodriguez Dominguez, Muhammad Shahzad, Xia Hong
arXiv AI
Sep 11

Teacher Geometry Shapes Learnability in Teacher-Student Networks

The paper investigates how the geometry of teacher neural networks affects the learnability of student networks in teacher‑student setups. By formalizing learnability as the success rate of reaching the global minimum, the authors identify two teacher distributions—one maximizing node dissimilarity (easy) and one minimizing it (hard)—that lead to markedly different success rates across various settings and activation functions. They analyze the loss landscape of small networks, revealing two types of suboptimal local minima (out‑of‑bounds and interior) whose attraction regions depend on teacher structure, and demonstrate that adjusting learning rates for the readout layer and inner biases can improve success rates. whyItMatters:"The study highlights that teacher geometry, often overlooked, plays a crucial role in determining how effectively a student network can learn, offering guidance for designing more realistic teacher‑student experiments."

By Kai J. Sandbrink, Flavio Martinelli, Alexander van Meegen, Wulfram Gerstner, Johanni Brea
arXiv Computer Vision
Sep 3

Breaking the Geometric Bottleneck: Contrastive Expansion in Asymmetric Cross-Modal Distillation

The paper investigates how knowledge distillation from Vision Transformers to smaller CNNs can cause dimensional collapse in the student’s representation space. Using SVD and Shannon entropy, the authors show that cosine‑based distillation leads to a drastic reduction in effective rank, while adding an InfoNCE objective can double the rank but harms downstream accuracy due to signal dilution. They further demonstrate that a label‑aware contrastive objective (Supervised Contrastive distillation) can maintain or improve accuracy without unnecessary rank expansion, indicating that effective rank alone is not a reliable indicator of representation quality.

By Kabir Thayani
arXiv AI
Jul 22

Soft-TransFormers for Continual Learning

arXiv:2411. 16073v4 Announce Type: replace-cross Abstract: Inspired by the Well-initialized Lottery Ticket Hypothesis (WLTH), we introduce Soft-TransFormers (Soft-TF), a continual learning framework that adapts a frozen pre-trained Transformer through task-specific soft subnetworks: real-valued multiplicative masks over the query, key, value, and output projections of selected self-attention layers.

By Haeyong Kang, Chang D. Yoo