arXiv AI

Teacher Geometry Shapes Learnability in Teacher-Student Networks

The paper investigates how the geometry of teacher neural networks affects the learnability of student networks in teacher‑student setups. By formalizing learnability as the success rate of reaching the global minimum, the authors identify two teacher distributions—one maximizing node dissimilarity (easy) and one minimizing it (hard)—that lead to markedly different success rates across various settings and activation functions. They analyze the loss landscape of small networks, revealing two types of suboptimal local minima (out‑of‑bounds and interior) whose attraction regions depend on teacher structure, and demonstrate that adjusting learning rates for the readout layer and inner biases can improve success rates. whyItMatters:"The study highlights that teacher geometry, often overlooked, plays a crucial role in determining how effectively a student network can learn, offering guidance for designing more realistic teacher‑student experiments."

arXiv Machine Learning
Sep 7

Coarse-Graining Hidden Representations: Unsupervised Neuron Selection via Mapping Entropy

The paper introduces an unsupervised method for selecting essential neurons in overparameterized neural networks by minimizing mapping entropy (ME), a metric that quantifies the loss of discriminatory power when neurons are discarded. ME-based selection relies solely on hidden-activation statistics and, in experiments, identifies minimal teacher-consistent representations in teacher‑student networks and coherent functional-class mappings in a non‑linear Gaussian process task. Subnetworks chosen by ME outperform random subsets of the same size, especially under strong compression, on both a Gaussian process task and translation‑augmented MNIST.

By Margherita Mele, Andrea Castagna, Roberto Menichetti, Raffaello Potestio, Alessandro Ingrosso