arXiv AI By Jiaheng Chen, Jiaxing Li, Yucheng Xiao, Xinyong Cai, Juncheng Bu, Lan Yu, Tinghe Zhang

Effective Does Not Mean Useful: Conditional Functional Substitutability for Redundancy and Scaling in Transformers

Read the original on arXiv AI →

The paper introduces Conditional Functional Substitutability (CFS) as a new way to measure redundancy in Transformers by examining when intermediate states produce similar downstream responses. CFS uncovers functional relationships and potential reductions that traditional importance- or similarity-based metrics miss, revealing systematic reorganization as models scale. Experiments across modalities and Transformer families show that performance gains do not always align with increased substitutability, and that models with more independent functional structure perform better, offering a functional explanation for diminishing returns and enabling more efficient dynamic computation.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Machine Learning
Sep 2

Performance-Efficiency Tradeoffs in Transformers: An Approximation Theory Perspective

The paper examines how to allocate attention heads and head dimensions across Transformer layers to balance expressivity and efficiency. It provides a mathematical analysis of early layers’ role in information extraction and characterizes the trade‑off between head count and dimension under a fixed parameter budget. The authors prove a saturation effect of softmax activations, showing that increasing head dimensions yields diminishing returns, especially for long sequences, and propose strategies for efficient parameter allocation across layers.

By Ruoxi Yu, Haotian Jiang, Jingpu Cheng, Penghao Yu, Qianxiao Li, Zhong Li
arXiv Machine Learning
Sep 22

The Ups and Downs of Backprop Weights

The paper discusses how backpropagation enables deep learning but does not inherently organize parameters for reusable functional components, leading to weight entanglement where overlapping parameter sets hinder independent modification. It introduces weight operators—parameterized modules that can be composed at inference—to address this, proposing a two-stage learning process that first infers operator composition and then updates only the selected operators. Vector Networks (VNs) are presented as an implementation that couples operator selection to local error-driven updates, demonstrating that learned operators can be recombined in unseen ways while keeping updates confined to the relevant parameter sets.

By Giuseppe Chindemi, Benjamin F. Grewe