arXiv:2511.09239v2 Announce Type: replace
Abstract: Deep neural networks typically learn spatially entangled representations that conflate discriminative foreground features with spurious background...
By Kaixiang Shu
arXiv:2606. 08156v1 Announce Type: cross Abstract: Vision Transformers (ViTs) achieve strong performance but suffer from high computational costs due to quadratic self-attention complexity.
By Kyumin Choi, Ikbeom Jang
arXiv:2605.12491v2 Announce Type: replace
Abstract: Vision Transformers (ViTs) learn rich visual-semantic representations through all-to-all self-attention among patch tokens. However, this design im...
By Alan Z. Song, Yinjie Chen, Mu Nan, Deva Ramanan, Michael J. Tarr, Andrew F. Luo
arXiv:2609.10387v1 Announce Type: new
Abstract: Deformable convolution networks have recently become popular for many computer vision tasks, especially for semantic segmentation, because of their exc...
By Yixiao Li, Xiaoyuan Yang, Jin Jiang, Minghao Zou, Guanghui Yue, Baoquan Zhao, Jun Liu, Wei Zhou
arXiv:2508.03351v3 Announce Type: replace-cross
Abstract: Large language models (LLMs) have demonstrated remarkable capabilities across diverse language tasks, motivating their extension to vision-la...
By Yufei Xue, Yushi Huang, Lunjie Zhu, Jiawei Shao, Jun Zhang
arXiv:2607. 08605v1 Announce Type: cross Abstract: Sparse autoencoders (SAEs) have emerged as a promising technique for mechanistic interpretability by learning a set of sparse latent features in large models, each of which encodes a distinct concept.
By Weiduo Liao, Yunqiao Yang, Ying Wei