arXiv AI

PixCon: Clean-Positive Contrastive Learning for Foundation-Model Semi-Supervised Segmentation

arXiv:2607. 03068v1 Announce Type: cross Abstract: Semi-supervised semantic segmentation (SSSS) has long turned on one question, which pseudo-labels to trust, and answered it with ever more careful confidence filtering.

arXiv Machine Learning
Aug 11

Contrastive Mask Fidelity: Reference-Free Auditing of Ground-Truth Masks in Remote Sensing Semantic Segmentation

arXiv:2608. 09101v1 Announce Type: cross Abstract: Semantic segmentation models are trained and evaluated against human-drawn masks, yet remote-sensing annotations are often coarse, incomplete, or misaligned; high overlap scores may then reflect agreement with imperfect labels rather than faithfulness to the image, creating an evaluation paradox.

By Shuaishuai Cao, Shuwei Peng, Meng Tang, Min Huang, Youjin Wang, Jie Chen, Jing Ouyang, Zhiwei Zhai
arXiv AI
Jun 16

ActiveSAM: Image-Conditional Class Pruning for Fast and Accurate Open-Vocabulary Segmentation

arXiv:2606. 16996v1 Announce Type: cross Abstract: Segment Anything Model 3 (SAM 3) provides a strong frozen backbone for concept-prompted segmentation, but applying it directly to open-vocabulary semantic segmentation (OVSS) is inefficient: full-resolution decoding is typically run over the entire dataset vocabulary, whereas each image contains only a small active subset of classes.

By Tran Dinh Tien, Zhiqiang Shen
arXiv Computer Vision
Sep 18

Distance to Class Prototypes: Active Learning for Object Detection

The paper introduces a new active learning signal for object detection that relies on a supervised contrastive term added to the training objective. This term shapes an embedding space where distance reflects class membership, allowing an unlabeled detection to be scored by its distance from the predicted category’s region weighted by confidence—all from a single forward pass of one network. Experiments on PASCAL VOC and MS‑COCO show that this criterion outperforms the standard posterior and remains competitive with ensemble‑based methods while incurring only a modest 8.3% increase in parameters.

By Licheng Zhang, Zheng Gong
arXiv AI
Jul 29

Why Does Grounding Hurt Medical VQA? Benchmarking, Diagnosis, and Fine-Tuning of Vision-Language Models

arXiv:2604. 27720v2 Announce Type: replace Abstract: Vision-language models (VLMs) are increasingly applied to medical visual question answering (Med-VQA), yet whether they can \emph{localize} the evidence behind their answers---a prerequisite for clinical auditability---is poorly characterized.

By Xupeng Chen, Binbin Shi, Chenqian Le, Qifu Yin, Lang Lin, Haowei Ni, Ran Gong, Panfeng Li
arXiv AI
Aug 6

Adversarially Robust Abductive Fusion of Pre-trained Transformer-based Perception Models

arXiv:2608. 04190v1 Announce Type: new Abstract: Deploying pre-trained perception models in novel environments degrades their accuracy under distributional shift, and assembling them alone does not recover it: combiners such as majority voting trade recall for precision and are brittle to coordinated failures.

By Mario Leiva, Yue Ma, Qinru Qiu, Gerardo Simari, Paulo Shakarian