arXiv AI

Balanced Prompt Adaptation against Entropy-Induced Collapse for Test-Time Binary Segmentation

arXiv Machine Learning
Aug 27

Fairness-Aware Test-Time Prompt Tuning

The paper introduces FairTPT, a fairness-aware test‑time prompt tuning method for vision‑language models like CLIP. It jointly minimizes target marginal entropy while maximizing spurious marginal entropy to reduce bias under subpopulation shifts. Experiments show that standard episodic test‑time adaptation can worsen disparities, but FairTPT outperforms existing debiasing methods while preserving overall performance.

By Yoann Launay, Parameswaran Kamalaruban, Tom Kempton, Stuart Burrell, David Sutton
arXiv Machine Learning
Sep 1

Towards Continual Test-Time Adaptation of Vision-Language Models in Open-Vocabulary Semantic Segmentation

arXiv:2608.29923v1 Announce Type: cross Abstract: Open-vocabulary semantic segmentation (OVSS) relies on vision-language alignment to recognize arbitrary text-defined categories, yet this alignment i...

By Chandler Timm C. Doloriel, Yunbei Zhang, Sarthak Kumar Maharana, Muhammad Salman Siddiqui, Tor Kristian Stevik, Fadi Al Machot, Kristian Hovde Liland, Habib Ullah
arXiv Computer Vision
Aug 24

ES-VP : Energy-Shaped Dynamic Visual Prompting for Efficient Model Adaptation

ES‑VP introduces Energy‑Shaped Visual Prompting, a method that generates image‑specific prompts through low‑rank initialization and energy‑guided dynamic adaptation. It achieves higher performance than existing single‑prompt and diverse‑prompt approaches while using far fewer parameters. Experiments on five architectures and fifteen datasets show consistent superiority, including a 2.6% accuracy gain over DAM‑VP on CLIP with 590× fewer prompt parameters.

By Can Jin, Ying Li, Jingchen Sun, Hongwu Peng, Jiahui Zhao, Yang Zhou, Lei Li, Dimitris N. Metaxas
arXiv AI
Jun 16

Mitigating Object Hallucinations in LVLMs via Attention Imbalance Rectification

arXiv:2603. 24058v2 Announce Type: replace-cross Abstract: Object hallucination in Large Vision-Language Models (LVLMs) severely compromises their reliability in real-world applications, posing a critical barrier to their deployment in high-stakes scenarios such as autonomous driving and medical image analysis.

By Han Sun, Qin Li, Peixin Wang, Min Zhang
arXiv AI
Jul 17

Parameter-efficient Prompt Tuning of Vision Foundation Model With Adaptive Focal Loss for Interpretable MCI Screening

arXiv:2607. 15047v1 Announce Type: cross Abstract: Mild Cognitive Impairment is a critical early stage of cognitive decline that frequently precedes Alzheimer's disease, yet its automated detection from neuropsychological drawing tests remains fundamentally constrained by data scarcity, class imbalance, and diagnostic ambiguity near clinical boundaries.

By Javad Khoramdel, Farhad Hoseyni, Amirhossein Nikoofard