arXiv Machine Learning By Yuxuan Gu, Wuyang Zhou, Huijun Xing, Danilo Mandic

TEMPER: Tensorized Efficient Manifold-constrained Parameterization for Expressive Residual Routing

Read the original on arXiv Machine Learning →

arXiv:2608. 07851v1 Announce Type: new Abstract: Residual connections rely on a static residual pathway, and are essential for training deep neural networks.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv AI
6d ago

Spectral-Sphere-Constrained Hyper-Connections

The paper introduces Spectral‑Sphere‑Constrained Hyper‑Connections (s²HC), a new method for controlling the residual matrices used in Hyper‑Connections (HC). Unlike previous doubly stochastic constraints that caused identity degeneration, expressivity bottlenecks, and parameterization inefficiencies, s²HC confines these matrices to a spectral norm sphere, restoring flexibility over subdominant spectra and eliminating unstable Sinkhorn‑Knopp iterations. This approach preserves training stability while allowing expressive, non‑degenerate residual matrices.

By Zhaoyi Liu, Haichuan Zhang, Ang Li
arXiv Machine Learning
Sep 14

Dead Weights, Live Signals: Feedforward Graphs of Frozen Language Models

The paper introduces a feedforward graph architecture that uses several frozen large language models as computational nodes connected through a shared continuous latent space via learned linear projections. By jointly optimizing projection matrices through backpropagation, the system combines the representations of three small frozen models with two larger ones, culminating in a lightweight cross‑attention output node. With only 17.6 M trainable parameters, the architecture attains state‑of‑the‑art results on ARC‑Challenge, OpenBookQA, and MMLU, surpassing both individual constituent models and parameter‑matched learned classifiers.

By Marcus Armstrong, Navid Ayoobi, Arjun Mukherjee
arXiv Machine Learning
Aug 27

GRIP: Algorithm-Agnostic Machine Unlearning for Mixture-of-Experts via Geometric Router Constraints

The paper introduces GRIP, an algorithm‑agnostic framework for machine unlearning in Mixture‑of‑Experts large language models. GRIP enforces hard geometric constraints on router updates, projecting gradient changes into the null space of the retain set’s routing matrix to prevent routing manipulation. Two variants—training‑time stochastic projection and post‑training analytical correction—show significant improvements in routing stability, retain accuracy, and resistance to white‑box adversarial recovery across two MoE models.

By Andy Zhu, Rongzhe Wei, Yupu Gu, Pan Li
arXiv Machine Learning
Jun 10

Rank Collapse, Fixed Points, and the Renormalization Group Structure of MLP Residual Networks

arXiv:2606. 10324v1 Announce Type: new Abstract: The analogy between deep neural network forward passes and renormalization group (RG) flows has been repeatedly noted in the literature, but existing treatments remain qualitative: depth is described as a coarse-graining scale, attention is likened to a partition function, and representations are said to flow toward fixed points.

By Parviz Haggi-Mani, Irina Rish