This paper demonstrate that whether masking-based token pruning helps or hurts worst-group robustness can be predicted before deployment, without labels or fine-tuning. A systematic study of semantic...
arXiv:2606. 04857v1 Announce Type: new Abstract: Standard IMVC evaluation retrains separate models for different missing-data configurations.
By Haolu Liu, Xiyue Wang, Xuanting Xie, Liangjian Wen, Zhao Kang
arXiv:2503. 09399v4 Announce Type: replace-cross Abstract: Large-scale image classification datasets exhibit strong compositional biases: objects tend to be centered, appear at characteristic scales, and co-occur with class-specific context.
By Tobias Christian Nauen, Brian Moser, Federico Raue, Stanislav Frolov, Andreas Dengel
arXiv:2609.39134v1 Announce Type: new
Abstract: Visual-token compression improves the efficiency of large vision-language models, but can expose failures that full-token evaluation misses. We study a...
By Shilinlu Yan, Bowen Chen, Yuechen Zhang, Zhenhong Zhou, Li Sun, Sen Su
arXiv:2608. 19285v1 Announce Type: cross Abstract: Recent Visual-Language Models (VLMs) have enhanced the capabilities of pre-trained LLMs by adding vision tokens alongside text, with approaches like LLaVA showing impressive results.
By Baptiste Rossigneux, Inna Kucher, Vincent Lorrain, Emmanuel Casseau
arXiv:2610. 01578v1 Announce Type: new Abstract: Why can masked prediction learn useful representations that unmasked reconstruction misses?
By Jorge Medina Moreira, Lorenzo Bardone, Lenka Zdeborov\'a
arXiv:2609.31558v1 Announce Type: new
Abstract: Contrastive Language--Image Pretraining (CLIP) has emerged as a dominant vision backbone due to its strong transferability and zero-shot capabilities....
By Ahmed Abdelnaby, Mohamed Elmahallawy
arXiv:2608. 09101v1 Announce Type: cross Abstract: Semantic segmentation models are trained and evaluated against human-drawn masks, yet remote-sensing annotations are often coarse, incomplete, or misaligned; high overlap scores may then reflect agreement with imperfect labels rather than faithfulness to the image, creating an evaluation paradox.
By Shuaishuai Cao, Shuwei Peng, Meng Tang, Min Huang, Youjin Wang, Jie Chen, Jing Ouyang, Zhiwei Zhai
SinkPruner is a training‑free framework that prunes visual tokens for multimodal large language models by first removing high‑norm redundant tokens with a visual sanitizer and then selectively keeping tokens that align with the text query using a text‑guided pruner. The coarse‑to‑fine design reduces attention sink and dispersion, enabling an 89% token reduction while preserving 96.5% of LLaVA‑1.5’s performance and 91.8% of Qwen2.5‑VL’s performance across twelve image‑language and four video‑language benchmarks. The visual sanitizer also improves existing pruning methods, showing strong transferability.
By Shiyu Li, Zi-Yuan Hu, Shijia Huang, Yanyang Li, Yiwu Zhong, Liwei Wang
arXiv:2608. 06901v1 Announce Type: cross Abstract: Vision-language models (VLMs) have achieved remarkable generalization across diverse multimodal tasks through large-scale pre-training, yet their rapidly increasing computational and memory requirements pose significant challenges for deployment in constrained environments.
By Minseok Kang, Hyunwoo Kim, Chanyoung Kim, Minwoo Kim, Jaekoo Lee, Dahuin Jung
arXiv:2608.23921v1 Announce Type: new
Abstract: Recent Vision-Language Models encode high-resolution images into long visual token sequences, incurring prohibitive prefill costs. To compress them, ex...
By Yuanhao Sun, Huawei Ji, Yuan Jin, Cheng Deng, Luoyi Fu, Xinbing Wang
arXiv:2608. 13711v1 Announce Type: cross Abstract: Computer-aided detection (CADe) systems for colonoscopy promise to reduce clinical miss rates, yet reliable real-world deployment remains elusive.
By Sebastian Doerrich, Andreas Franz Schwab, Francesco Di Salvo, Shyam Nandan Rai, Hanh Huyen My Nguyen, Christian Ledig