arXiv:2607. 00620v1 Announce Type: cross Abstract: Generalized Category Discovery (GCD) aims to recognize known classes while autonomously discovering novel ones in open-world settings.
By Boyang Dai, Chaoqi Chen, Yizhou Yu
arXiv:2605. 30170v2 Announce Type: replace-cross Abstract: While Large Vision-Language Models (VLMs) excel at interpolation, they suffer catastrophic failures in systematic generalization, most notably in visual counting.
By Xingzhou Pang, Yifan Hou, Junling Wang, Mrinmaya Sachan
arXiv:2606. 16193v1 Announce Type: cross Abstract: Multimodal Large Language Models (MLLMs) have demonstrated strong performance on vision-language tasks, yet their internal visual representations remain difficult to interpret.
By Yusong Zhao, Hengyi Wang, Tanuja Ganu, Akshay Nambi, Hao Wang
arXiv:2606. 03569v1 Announce Type: cross Abstract: Vision-Language Models (VLMs) have demonstrated remarkable capabilities but suffer from significant computational overhead during inference.
By Jiahui Wang, Kai Zhang, Mai Han, Huanghe Zhang
arXiv:2606. 00082v1 Announce Type: cross Abstract: Explainability of deep learning algorithms is critical for computer-vision applications with high-stake decisions.
By Cl\'ement B\'enard, Manon Arfib, Christophe Labreuche, Victor Qu\'etu
Vision-Language Models (VLMs) have demonstrated remarkable capabilities but suffer from significant computational overhead during inference. While visual token pruning offers a promising solution, existing methods predominantly rely on initial attention scores.
arXiv:2511. 10260v2 Announce Type: replace-cross Abstract: Fine-Grained Visual Classification (FGVC) remains a challenging task due to subtle inter-class differences and large intra-class variations.
By Yongji Zhang, Siqi Li, Kuiyang Huang, Yue Gao, Yu Jiang
arXiv:2606. 28399v1 Announce Type: cross Abstract: The structure of human visual representations underpins our capacity for adaptive behaviour.
By Can Demircan, Marcel Binz, Alireza Modirshanechi, Eric Schulz
arXiv:2607. 04548v1 Announce Type: cross Abstract: Novel category discovery aims to identify unseen classes from unlabeled data by transferring knowledge from labeled categories, but most existing methods perform discovery in opaque latent feature spaces.
By Ifrat Ikhtear Uddin, Yang Zhou, KC Santosh, Longwei Wang
arXiv:2608. 11197v1 Announce Type: new Abstract: Shani et al.
By Nikolai Bolik, Lennart St\"opler, Artur Andrzejak
arXiv:2604. 07753v2 Announce Type: replace-cross Abstract: Empowering Large Multimodal Models (LMMs) with image generation often leads to catastrophic forgetting in understanding tasks due to severe gradient conflicts.
By Xiangyue Liu, Zijian Zhang, Miles Yang, Zhao Zhong, Liefeng Bo, Ping Tan
arXiv:2606. 02172v1 Announce Type: new Abstract: Learning discriminative visual representations from distributed, heterogeneous data is a fundamental challenge in Federated Learning (FL).
By Mario Casado-Diez, Alejandro Dopico-Castro, Ver\'onica Bol\'on-Canedo, Bertha Guijarro-Berdi\~nas