Learned Queries and Keys Are All You Need: Replacing the Value Projection with Structured Transforms
Read the original on arXiv Machine Learning →The Flow has not summarised this story yet — read it at arXiv Machine Learning.
The Flow has not summarised this story yet — read it at arXiv Machine Learning.
arXiv:2410.14874v3 Announce Type: replace Abstract: Multi-Head Self-Attention (MHSA) is the cornerstone of Vision Transformers, allowing models to capture diverse feature representations by projectin...
The paper introduces Transformer-Within-Transformer (TWT), a post‑hoc technique that merges contiguous redundant layers in Vision Transformers into a single surrogate layer. By doing so, TWT cuts both parameter count and inference compute while maintaining competitive performance on natural image tasks with only half the depth. In histopathology applications, TWT not only matches but sometimes surpasses the baseline model’s performance.
The paper examines how to allocate attention heads and head dimensions across Transformer layers to balance expressivity and efficiency. It provides a mathematical analysis of early layers’ role in information extraction and characterizes the trade‑off between head count and dimension under a fixed parameter budget. The authors prove a saturation effect of softmax activations, showing that increasing head dimensions yields diminishing returns, especially for long sequences, and propose strategies for efficient parameter allocation across layers.
arXiv:2602. 06883v3 Announce Type: replace Abstract: The smoothness of the transformer architecture has been extensively studied in the context of generalization, training stability, and adversarial robustness.
arXiv:2607. 03612v1 Announce Type: cross Abstract: Feed-forward 3D reconstruction (F3R) transformers have recently achieved remarkable success.
arXiv:2606. 13276v1 Announce Type: cross Abstract: Weight-space geometry plays a central role in neural network optimization, yet manifold constraints are often applied uniformly across all weight matrices.