arXiv:2607. 02612v1 Announce Type: cross Abstract: Vision Transformers achieve strong image classification accuracy but process all image regions with nearly the same computation, even when many regions are redundant or uninformative.
By Aravind Pradeep, Samira Nazari, Mahdi Taheri, Christian Herglotz
arXiv:2605.08371v2 Announce Type: replace
Abstract: Multi-view geometry transformers are feed-forward 3D foundation models that jointly predict depth maps, point maps, and camera poses for N images i...
By Haotang Li, Zhenyu Qi, Shaohan Henry Wang, Kebin Peng, Zi Wang, Qing Guo, Sen He, Huanrui Yang
The paper introduces Quantizer‑Aligned Recalibration (QuAR), a single‑pass test‑time adaptation technique for quantized vision transformers that does not require backpropagation or parameter updates. QuAR recalibrates activations at the input of frozen quantizers by aligning per‑channel statistics with the source calibration, thereby correcting the distorted code distribution caused by distribution shift. On ImageNet‑C, QuAR outperforms state‑of‑the‑art backprop‑free methods across 3‑, 4‑, 6‑, and 8‑bit precisions, achieving higher accuracy, lower latency, and minimal memory overhead while maintaining performance across diverse shift scenarios.
By Hyeongheon Cha, Young D. Kwon, Sung-Ju Lee
arXiv:2607. 03784v1 Announce Type: cross Abstract: While prior studies have successfully compressed vision Transformers (ViTs) through various pruning techniques, most have concentrated on width pruning to achieve significant reductions in model size.
By Zhenfeng Su, Kang Zhao, Han Bao, Tao Yuan, Zhongzhe Hu, Xianzhi Yu, Wenxuan Wang
arXiv:2602. 20114v2 Announce Type: replace-cross Abstract: Machine unlearning (MU) refers to the post-training capability to remove (the influence of) training examples that are incorrect, biased, or leak sensitive/private information.
By Kairan Zhao, Iurie Luca, Peter Triantafillou
FORGE is a forward‑only test‑time adaptation technique designed for integer‑only vision models running on microcontrollers. It restores batch‑normalization statistics after BN folding by re‑normalizing each convolution’s per‑channel output using only forward‑pass estimates, enabling adaptation on deployed, folded integer models. The method achieves accuracy gains comparable to gradient‑based TENT, requires adapting only a few layers, works with single‑sample streaming, and has been validated on an ESP32‑S3 with minimal energy and latency overhead.
By Muhammad Rehan, Haider Ali, Muhammad Ali Munir, Moaz Amjad
arXiv:2608.28706v1 Announce Type: new
Abstract: ViT detectors fix a uniform token grid before any learned stage. A native-resolution aerial detector must then choose between resolving few-pixel objec...
By Khayrul Islam
arXiv:2603. 00198v2 Announce Type: replace-cross Abstract: Token reduction accelerates long-video vision--language models (VLMs), but existing methods target Transformers, where reduction is treated as token pruning.
By Jindong Jiang, Amala Sanjay Deshmukh, Kateryna Chumachenko, Karan Sapra, Zhiding Yu, Guilin Liu, Andrew Tao, Pavlo Molchanov, Jan Kautz, Wonmin Byeon
arXiv:2603. 12478v2 Announce Type: replace-cross Abstract: Multimodal instruction tuning is often compute-inefficient because training budgets are spread across large mixed image-video pools whose utility is highly uneven.
By Rujie Wu, Haozhe Zhao, Hai Ci, Yizhou Wang
arXiv:2505. 15441v5 Announce Type: replace-cross Abstract: Natural images exhibit strong geometric regularities: local structures, such as edges, corners, and textures, appear in many orientations and mirror configurations.
By David Nordstr\"om, Johan Edstedt, Fredrik Kahl, Georg B\"okman
Fewer visual tokens do not guarantee lower end-to-end latency. We evaluate break-even with a reproducible protocol that accounts for decision overhead, shared work, and the operators each policy can avoid.
arXiv:2608. 13141v1 Announce Type: cross Abstract: Vision Transformers (ViTs) demonstrate exceptional performance in computer vision but suffer from large parameter counts and quadratic computational complexity, severely limiting their deployment on resource-constrained edge hardware.
By Junseo Kim, Uraz Odyurt, Amirreza Yousefzadeh