arXiv:2609.26167v1 Announce Type: new
Abstract: Activation-energy pruning -- removing weights whose product of magnitude and cumulative pre-synaptic spike count falls below a threshold -- was establi...
By Joseph Bingham
The paper investigates how to choose the best quantized model from a family of compressed versions when target labels are scarce or unavailable. It finds that a simple rule based on minimum teacher distortion consistently selects the same eight‑bit, per‑channel, unclipped configuration, though this does not minimize empirical target cross‑entropy. The study also shows that confidence‑based estimators perform poorly in overconfident regimes, while output‑distribution estimators can outperform the teacher in some architectures, and that combining distortion with a supervised term can improve selection. Across 134 candidate families, teacher‑anchored selection reduces mean regret with very few labels, though the benefit diminishes after about 25 labels.
By Alejandro Rodriguez Dominguez, Muhammad Shahzad, Xia Hong
arXiv:2609.24379v1 Announce Type: cross
Abstract: Mechanistic interpretability of vision transformers seeks to decompose model computation into human-readable units, but learned representations entan...
By Gautam Ranka, Shubham Santosh Pandere, Aiden Dsouza
The paper investigates how the geometry of teacher neural networks affects the learnability of student networks in teacher‑student setups. By formalizing learnability as the success rate of reaching the global minimum, the authors identify two teacher distributions—one maximizing node dissimilarity (easy) and one minimizing it (hard)—that lead to markedly different success rates across various settings and activation functions. They analyze the loss landscape of small networks, revealing two types of suboptimal local minima (out‑of‑bounds and interior) whose attraction regions depend on teacher structure, and demonstrate that adjusting learning rates for the readout layer and inner biases can improve success rates.
whyItMatters:"The study highlights that teacher geometry, often overlooked, plays a crucial role in determining how effectively a student network can learn, offering guidance for designing more realistic teacher‑student experiments."
By Kai J. Sandbrink, Flavio Martinelli, Alexander van Meegen, Wulfram Gerstner, Johanni Brea
Mechanistic interpretability of vision transformers seeks to decompose model computation into human-readable units, but learned representations entangle many concepts in each neuron. Feature superposi...
arXiv:2606. 07414v1 Announce Type: new Abstract: Sparsity allows scaling model parameters without proportionally increasing computational cost.
By Simon Schug
arXiv:2608. 06766v1 Announce Type: cross Abstract: Training changes a network's predictions while allocating task-relevant structure across its internal units.
By Tongxi Wang
arXiv:2512. 11000v2 Announce Type: replace-cross Abstract: Representations pervade our daily experience, from letters representing sounds to bit strings encoding digital files.
By Francesco L\"assig
The paper investigates how knowledge distillation from Vision Transformers to smaller CNNs can cause dimensional collapse in the student’s representation space. Using SVD and Shannon entropy, the authors show that cosine‑based distillation leads to a drastic reduction in effective rank, while adding an InfoNCE objective can double the rank but harms downstream accuracy due to signal dilution. They further demonstrate that a label‑aware contrastive objective (Supervised Contrastive distillation) can maintain or improve accuracy without unnecessary rank expansion, indicating that effective rank alone is not a reliable indicator of representation quality.
By Kabir Thayani
arXiv:2411. 16073v4 Announce Type: replace-cross Abstract: Inspired by the Well-initialized Lottery Ticket Hypothesis (WLTH), we introduce Soft-TransFormers (Soft-TF), a continual learning framework that adapts a frozen pre-trained Transformer through task-specific soft subnetworks: real-valued multiplicative masks over the query, key, value, and output projections of selected self-attention layers.
By Haeyong Kang, Chang D. Yoo
arXiv:2510. 24616v4 Announce Type: replace-cross Abstract: For four decades statistical physics has been providing a framework to analyse neural networks.
By Jean Barbier, Francesco Camilli, Minh-Toan Nguyen, Mauro Pastore, Rudy Skerk
arXiv:2606. 05863v1 Announce Type: new Abstract: Grokking suggests that fitting the training data and learning a simple underlying rule may occur on different time scales.
By Hu Tan, Kuo Gai, Shihua Zhang