Methods operating on Vision Transformer (ViT) feature spaces typically rely on Euclidean distance or cosine similarity. This assumes that every direction is equally meaningful, but there is no reason...
arXiv:2606. 02765v1 Announce Type: cross Abstract: Model dimension ($d_{model}$) is a fundamental hyperparameter in transformer language models, yet its role in setting the geometric limits of feature representation remains under-explored.
By Alexander Guha
arXiv:2608. 01283v1 Announce Type: new Abstract: All Transformer-based large language models compute attention via the Euclidean inner product, an architectural choice that Dong et al.
By Sen Song
arXiv:2609.37717v1 Announce Type: new
Abstract: Decoder-only transformers are trained only through a terminal next-token prediction loss, yet this loss constrains every intermediate hidden state thro...
By Timur Mudarisov, Mikhail Burtsev, Tatiana Petrova, Radu State
Cut‑ViT introduces a task‑specific pruning pipeline for visual foundation models that uses gram anchoring matrices and subspace decomposition to align feature representations between native and pruned DINOv3 models. The method incorporates basis‑agnostic and residual constraints to preserve robustness across spatial and channel dimensions, and employs spectral entropy adaptation to tailor the pruning objective to downstream tasks. Experiments demonstrate that Cut‑ViT achieves state‑of‑the‑art performance on six tasks across nine datasets while reducing pruning time to about one minute on a single A100 GPU, using only 20.9% of the time and 45.5% of the GPU memory compared to prior methods.
By Jianjian Yin, Liulei Li, Tao Chen, Yi Chen, Yazhou Yao, Wenguan Wang
arXiv:2607. 02386v1 Announce Type: cross Abstract: While Vision Transformers have achieved remarkable success across computer vision and language applications, the geometric evolution of their internal representations throughout training remains insufficiently understood.
By Kaustubh Kapil, Kishor P. Upla
The paper introduces Successive Capacity Growth (SCG), a method for adaptively expanding Vision Transformer encoders in Joint-Embedding Predictive Architectures (JEPAs). SCG starts with a minimal encoder and incrementally increases width or depth based on a task‑agnostic test‑and‑verify mechanism, while a Sketched Isotropic Gaussian Regularizer (SIGReg) keeps learned semantic dimensions independent. Experiments on multi‑object dynamics and 2D navigation tasks show that SCG achieves up to 20.3% better prediction loss than fixed small baselines and 23% better than fixed large models, with far greater parameter efficiency and no false‑positive expansions.
By Frederik Berenz
The paper introduces Riemannian–Lorentz Parameter Fusion (RLPF), a method for merging a Vision Transformer and a state‑space model without gradient descent. RLPF aligns parameter groups by semantic role, projects them onto a common coordinate system, lifts selected coordinates to the Lorentz hyperboloid, computes a regularized geodesic barycenter, and decodes the result back into the two branches, with a learned gate combining their logits. The resulting fine‑tuned system achieves 82.37 % on CIFAR‑10, 75.04 % on Oxford‑IIIT Pet, and 78.58 % top‑1 accuracy on ImageNet‑1K, surpassing the best‑parent accuracies of 76.54 %, 71.42 %, and 76.42 % respectively.
By Badri N. Patro, Vijay S. Agneeswaran
arXiv:2608. 11661v1 Announce Type: cross Abstract: A multiplicative dual-encoder network computes a real-valued output for a pair of inputs as the inner product of their separate encodings.
By Zijian Zhao, Sen Li
The paper introduces Successive Capacity Growth (SCG), a method that starts with a minimal Vision Transformer encoder and incrementally expands its width or depth based on a task‑agnostic test‑and‑verify mechanism. SCG uses function‑preserving expansion and a Sketched Isotropic Gaussian Regularizer (SIGReg) to ensure independent semantic dimensions and prevent collapse. Experiments on multi‑object dynamics and 2D navigation tasks show that SCG achieves significant prediction loss reductions while being far more parameter‑efficient than fixed large models, with no false‑positive expansions and exact function preservation.
arXiv:2606. 01443v1 Announce Type: cross Abstract: A central difficulty in training Joint-Embedding Predictive Architectures (JEPAs) is preventing representation collapse.
By Triet M. Le
arXiv:2609.07937v1 Announce Type: cross
Abstract: Structured visual reasoning, such as image puzzles, demands fine-grained visual perception, an ability current Vision Language Models (VLMs) lack. VL...
By Harsha Patnala, Debopriyo Banerjee, Ayush Sunil Munot, Somak Aditya