arXiv Machine Learning By Giuseppe Chindemi, Benjamin F. Grewe

The Ups and Downs of Backprop Weights

Read the original on arXiv Machine Learning →

The paper discusses how backpropagation enables deep learning but does not inherently organize parameters for reusable functional components, leading to weight entanglement where overlapping parameter sets hinder independent modification. It introduces weight operators—parameterized modules that can be composed at inference—to address this, proposing a two-stage learning process that first infers operator composition and then updates only the selected operators. Vector Networks (VNs) are presented as an implementation that couples operator selection to local error-driven updates, demonstrating that learned operators can be recombined in unseen ways while keeping updates confined to the relevant parameter sets.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

Hugging Face Trending Papers
Jun 22

Sublinearly Structured Deep Neural Networks Achieve Feature Learning Consistency for Compositional Functions

Over the past decade, deep neural networks (DNNs) have achieved remarkable success on complex machine-learning tasks, yet the theoretical foundations of their performance remain incomplete. From a statistical viewpoint, a natural question is: can DNNs attain feature-learning and prediction consistency comparable to that of classical models?

arXiv Machine Learning
1d ago

CrossGMN: Graph Metanetworks for Cross-Architecture Weight-Space Transformations

CrossGMN introduces a graph metanetwork that processes a trained source network and an initialized target network simultaneously, enabling equivariant cross‑architecture weight‑space transformations. By preserving symmetry through cross‑network message passing, CrossGMN can refine target network initializations while remaining invariant to source permutations and equivariant to target permutations. Experiments demonstrate that CrossGMN accelerates knowledge distillation, transfers across datasets without retraining, and unifies compression from diverse source architectures into a common target architecture.

By Adir Dayan, Yam Eitan, Haggai Maron