arXiv Computer Vision
Sep 18

A Smaller Transformer in Your Transformer

The paper introduces Transformer-Within-Transformer (TWT), a post‑hoc technique that merges contiguous redundant layers in Vision Transformers into a single surrogate layer. By doing so, TWT cuts both parameter count and inference compute while maintaining competitive performance on natural image tasks with only half the depth. In histopathology applications, TWT not only matches but sometimes surpasses the baseline model’s performance.

By Dhananjay Tomar, Marius Aasan, Andreas Kleppe, Ad\'in Ram\'irez Rivera
arXiv Machine Learning
Sep 2

Performance-Efficiency Tradeoffs in Transformers: An Approximation Theory Perspective

The paper examines how to allocate attention heads and head dimensions across Transformer layers to balance expressivity and efficiency. It provides a mathematical analysis of early layers’ role in information extraction and characterizes the trade‑off between head count and dimension under a fixed parameter budget. The authors prove a saturation effect of softmax activations, showing that increasing head dimensions yields diminishing returns, especially for long sequences, and propose strategies for efficient parameter allocation across layers.

By Ruoxi Yu, Haotian Jiang, Jingpu Cheng, Penghao Yu, Qianxiao Li, Zhong Li