arXiv:2607. 28589v1 Announce Type: cross Abstract: Post-training quantization (PTQ) has emerged as an effective solution for deploying Vision Transformers (ViTs) on resource-constrained devices.
By Md. Mehrab Hossain Opi, Robiul Islam Ryad, Md. Umar Faruk
The paper introduces Transformer-Within-Transformer (TWT), a post‑hoc technique that merges contiguous redundant layers in Vision Transformers into a single surrogate layer. By doing so, TWT cuts both parameter count and inference compute while maintaining competitive performance on natural image tasks with only half the depth. In histopathology applications, TWT not only matches but sometimes surpasses the baseline model’s performance.
By Dhananjay Tomar, Marius Aasan, Andreas Kleppe, Ad\'in Ram\'irez Rivera
Visual Geometry Grounded Transformer (VGGT) recovers dense 3D scene structure from multi-view images in one forward pass, but quadratic cross-frame attention limits its scalability. Existing training-free accelerators reduce computation uniformly along one axis, missing layer heterogeneity.
The paper introduces a top‑down framework for accurately locating athletes in metric world coordinates using a single calibrated broadcast frame. It presents three main contributions: a Boundary‑Aware Adaptive Tiling method that expands tile boundaries to avoid splitting athletes across tiles, a specialized two‑keypoint estimator based on RTMPose‑X for pelvis and ground projection points, and a deterministic lift of 2D projections into 3D world coordinates via camera‑calibrated ray casting. The approach achieves a LocSim score of 97.44 and an mAP of 0.9128, surpassing the baseline by over 21 % on a public test set.
By Thanh-Khoi Nguyen, Hoang-Phuc Nguyen, Linh-Huynh, Minh-Triet Tran
GramLoop is a training‑free framework that enhances frozen DINOv3 dense‑prediction models under distribution shift by adding inference computation within the visual backbone. It replays a short transformer window and uses final‑layer cosine‑Gram consistency to control each replay, propagating proposals through the frozen suffix and accepting them via a patchwise gate. Across object detection and semantic segmentation tasks, GramLoop improves performance on all five shifted benchmarks, notably raising COCO‑O mAP by +0.252 and Effective Robustness by +0.250 while maintaining clean ADE20K accuracy.
By Yang Chen, Canyu Shen, Xinzhe Rao, Yuanyi Yan, Yunlu Chen, Meng Tang, Teng Long, Vincent Tao Hu
arXiv:2609.23733v1 Announce Type: new
Abstract: Feed-forward visual geometry models such as the Visual Geometry Grounded Transformer (VGGT) have recently enabled direct 3D reconstruction from multi-v...
By Abteen Arab, Guile Wu, Chengjie Huang, Dongfeng Bai