arXiv Machine Learning By O\u{g}uzhan Ercan

Phase Marginalization for Patch-Grid Instability in Vision Transformers

Read the original on arXiv Machine Learning →

arXiv:2606. 08132v1 Announce Type: cross Abstract: Vision Transformers operate on fixed patch grids, which can introduce phase-dependent instability for dense prediction: changing the patch partition can change the token evidence available to a pixel, especially near boundaries.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv Computer Vision
Sep 18

A Smaller Transformer in Your Transformer

The paper introduces Transformer-Within-Transformer (TWT), a post‑hoc technique that merges contiguous redundant layers in Vision Transformers into a single surrogate layer. By doing so, TWT cuts both parameter count and inference compute while maintaining competitive performance on natural image tasks with only half the depth. In histopathology applications, TWT not only matches but sometimes surpasses the baseline model’s performance.

By Dhananjay Tomar, Marius Aasan, Andreas Kleppe, Ad\'in Ram\'irez Rivera
arXiv Computer Vision
Sep 3

A Top-Down Framework for Metric-Scale Athlete Localization from Single Broadcast Frames

The paper introduces a top‑down framework for accurately locating athletes in metric world coordinates using a single calibrated broadcast frame. It presents three main contributions: a Boundary‑Aware Adaptive Tiling method that expands tile boundaries to avoid splitting athletes across tiles, a specialized two‑keypoint estimator based on RTMPose‑X for pelvis and ground projection points, and a deterministic lift of 2D projections into 3D world coordinates via camera‑calibrated ray casting. The approach achieves a LocSim score of 97.44 and an mAP of 0.9128, surpassing the baseline by over 21 % on a public test set.

By Thanh-Khoi Nguyen, Hoang-Phuc Nguyen, Linh-Huynh, Minh-Triet Tran
arXiv Computer Vision
Sep 1

GramLoop: Training-Free Gram-Gated Replay for Robust Dense Prediction

GramLoop is a training‑free framework that enhances frozen DINOv3 dense‑prediction models under distribution shift by adding inference computation within the visual backbone. It replays a short transformer window and uses final‑layer cosine‑Gram consistency to control each replay, propagating proposals through the frozen suffix and accepting them via a patchwise gate. Across object detection and semantic segmentation tasks, GramLoop improves performance on all five shifted benchmarks, notably raising COCO‑O mAP by +0.252 and Effective Robustness by +0.250 while maintaining clean ADE20K accuracy.

By Yang Chen, Canyu Shen, Xinzhe Rao, Yuanyi Yan, Yunlu Chen, Meng Tang, Teng Long, Vincent Tao Hu