arXiv Machine Learning By Joonas J\"arve, Halil Ibrahim Aysel, Tarun Khajuria, Meelis Kull

Feature Evolution and Migration during Vision Transformer Training

Read the original on arXiv Machine Learning →

arXiv:2608. 20134v1 Announce Type: cross Abstract: We present a novel view on feature evolution in Vision Transformers (ViTs) by visualizing the training process over two dimensions -- network depth (layer) and training time (epochs).

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv AI
4d ago

Geometric Similarity in VLM Low-Level Vision Representations

The paper introduces GeoSim, a four‑level framework for analyzing how vision‑language models (VLMs) represent low‑level vision tasks. It evaluates hidden‑layer representations across 24 tasks and two VLM paradigms—autoregressive models and diffusion transformers—using global similarity, local geometry, sparse feature decomposition, and topological verification. The study uncovers the organizing principles of low‑level visual representations and highlights their limitations in cross‑task and cross‑model agreement, offering an interpretability lens for assessing latent transferability and diagnosing model‑specific issues.

By Shao-Jun Xia, Huixin Zhang, Zhen Lei, Anlan Sun, Yuner Zhang, Xiaoyang Chen
arXiv Computer Vision
Sep 18

A Smaller Transformer in Your Transformer

The paper introduces Transformer-Within-Transformer (TWT), a post‑hoc technique that merges contiguous redundant layers in Vision Transformers into a single surrogate layer. By doing so, TWT cuts both parameter count and inference compute while maintaining competitive performance on natural image tasks with only half the depth. In histopathology applications, TWT not only matches but sometimes surpasses the baseline model’s performance.

By Dhananjay Tomar, Marius Aasan, Andreas Kleppe, Ad\'in Ram\'irez Rivera