arXiv:2606. 02339v1 Announce Type: new Abstract: Entropy minimization (EM) is the dominant objective for test-time adaptation, yet its failure mode, model collapse, remains poorly understood.
By Tim Nielen, Sameer Ambekar, Johannes Kiechle, Daniel M. Lang, Julia A. Schnabel
arXiv:2608.22996v1 Announce Type: new
Abstract: Vision-Language Models (VLMs) perform well on diverse vision-language tasks, but transformer-based visual encoders split images into fixed-resolution s...
By Yuanhao Sun, Huawei Ji, Jiaxin Ding, Luoyi Fu, Xinbing Wang
Vision-Language Models (VLMs) perform well on diverse vision-language tasks, but transformer-based visual encoders split images into fixed-resolution sub-images, compromising object integrity in light...
The paper introduces FairTPT, a fairness-aware test‑time prompt tuning method for vision‑language models like CLIP. It jointly minimizes target marginal entropy while maximizing spurious marginal entropy to reduce bias under subpopulation shifts. Experiments show that standard episodic test‑time adaptation can worsen disparities, but FairTPT outperforms existing debiasing methods while preserving overall performance.
By Yoann Launay, Parameswaran Kamalaruban, Tom Kempton, Stuart Burrell, David Sutton
arXiv:2503. 09399v4 Announce Type: replace-cross Abstract: Large-scale image classification datasets exhibit strong compositional biases: objects tend to be centered, appear at characteristic scales, and co-occur with class-specific context.
By Tobias Christian Nauen, Brian Moser, Federico Raue, Stanislav Frolov, Andreas Dengel
arXiv:2608.29923v1 Announce Type: cross
Abstract: Open-vocabulary semantic segmentation (OVSS) relies on vision-language alignment to recognize arbitrary text-defined categories, yet this alignment i...
By Chandler Timm C. Doloriel, Yunbei Zhang, Sarthak Kumar Maharana, Muhammad Salman Siddiqui, Tor Kristian Stevik, Fadi Al Machot, Kristian Hovde Liland, Habib Ullah
arXiv:2606. 07647v1 Announce Type: cross Abstract: Large vision language models (LVLMs) have made rapid advancements and are deployed across various applications, yet hallucinations remain a major challenge.
By Ruipeng Zhang, Zhihao Li, C. L. Philip Chen, Tong Zhang
ES‑VP introduces Energy‑Shaped Visual Prompting, a method that generates image‑specific prompts through low‑rank initialization and energy‑guided dynamic adaptation. It achieves higher performance than existing single‑prompt and diverse‑prompt approaches while using far fewer parameters. Experiments on five architectures and fifteen datasets show consistent superiority, including a 2.6% accuracy gain over DAM‑VP on CLIP with 590× fewer prompt parameters.
By Can Jin, Ying Li, Jingchen Sun, Hongwu Peng, Jiahui Zhao, Yang Zhou, Lei Li, Dimitris N. Metaxas
arXiv:2603. 24058v2 Announce Type: replace-cross Abstract: Object hallucination in Large Vision-Language Models (LVLMs) severely compromises their reliability in real-world applications, posing a critical barrier to their deployment in high-stakes scenarios such as autonomous driving and medical image analysis.
By Han Sun, Qin Li, Peixin Wang, Min Zhang
arXiv:2607. 15047v1 Announce Type: cross Abstract: Mild Cognitive Impairment is a critical early stage of cognitive decline that frequently precedes Alzheimer's disease, yet its automated detection from neuropsychological drawing tests remains fundamentally constrained by data scarcity, class imbalance, and diagnostic ambiguity near clinical boundaries.
By Javad Khoramdel, Farhad Hoseyni, Amirhossein Nikoofard
arXiv:2605. 03403v2 Announce Type: replace-cross Abstract: Group Relative Policy Optimization (GRPO) has recently shown strong performance in post-training large language models and vision-language models.
By Yujun Li, Hongyuan Zhang, Yuan Yuan
arXiv:2609.16656v1 Announce Type: new
Abstract: State space models (SSMs), particularly Mamba, have emerged as efficient alternatives to attention-based architectures and have been extended to vision...
By Jonghyeon Lim, Changhoon Yim