arXiv Computer Vision

Can Frozen Hyperspherical Features Guide the Selection of Pseudo Masks?

The paper introduces SphereTrust, a method that uses frozen self‑supervised hyperspherical features to evaluate and rank candidate masks produced by foundation segmenters like SAM. By measuring angular contrast, foreground coverage, and image‑frame contact, SphereTrust can select high‑quality masks in 0.55 s per image and outperforms existing baselines on multiple segmentation tasks. The selected masks are then used as priors to train student models, improving performance on several benchmark datasets.

arXiv AI
Aug 6

Adversarially Robust Abductive Fusion of Pre-trained Transformer-based Perception Models

arXiv:2608. 04190v1 Announce Type: new Abstract: Deploying pre-trained perception models in novel environments degrades their accuracy under distributional shift, and assembling them alone does not recover it: combiners such as majority voting trade recall for precision and are brittle to coordinated failures.

By Mario Leiva, Yue Ma, Qinru Qiu, Gerardo Simari, Paulo Shakarian
arXiv Computer Vision
Sep 18

Queries Knew More Than We Thought: Uncovering Latent Knowledge in Segmentation Models

The paper investigates an output‑selection bottleneck in frozen DETR‑family segmentation models, showing that many useful mask proposals are computed but never exposed. A lightweight selector called HYDRA, trained only on cached outputs, can recover up to +7.41 mIoU on ADE20k and COCO and +9.4 class‑macro prompt‑IoU on SAM 3 across eight domains by selectively choosing better candidates. The study demonstrates that evaluating segmenters should consider both exposed masks and the hidden candidates they suppress.

By Ignacio M. De la Jara, Cristian Rodriguez-Opazo, Damith Ranasinghe
arXiv Machine Learning
Sep 24

Feed the Panel Dimensions, Not Verdicts: Rubric-Decomposed Fusion of Vision-Language Aesthetic Judges

The paper investigates whether panels of vision‑language models (VLMs) can reliably judge image aesthetics. It shows that a panel of holistic judges rarely outperforms its best member, but when each model scores images on five rubric‑defined dimensions and these dimension scores are fused across model families, the panel consistently beats the best single VLM on two datasets (EVA and PARA). The study demonstrates that the value of a panel depends on the type of input it receives, and that dimension‑based fusion yields measurable gains at the cost of additional labeling and API usage.

By Amit Jadhav, Shaurya Beriwala, Beomjin Kim
arXiv Machine Learning
Aug 11

Contrastive Mask Fidelity: Reference-Free Auditing of Ground-Truth Masks in Remote Sensing Semantic Segmentation

arXiv:2608. 09101v1 Announce Type: cross Abstract: Semantic segmentation models are trained and evaluated against human-drawn masks, yet remote-sensing annotations are often coarse, incomplete, or misaligned; high overlap scores may then reflect agreement with imperfect labels rather than faithfulness to the image, creating an evaluation paradox.

By Shuaishuai Cao, Shuwei Peng, Meng Tang, Min Huang, Youjin Wang, Jie Chen, Jing Ouyang, Zhiwei Zhai
arXiv AI
Aug 18

Picking the Right Image to Classify: Reliable-Input Selection in Teledermatology

arXiv:2608. 16198v1 Announce Type: cross Abstract: Dermatology models face distribution shifts in teledermatology settings, where submitted images differ from the training data in lighting, angle, distance, focus, and framing.

By Fabian Gr\"oger, Marco Weishaupt, Philippe Gottfrois, Simone Lionetti, Linda Wermelinger, Nipun Ranasekara, Ludovic Amruthalingam, Alexander A. Navarini, Marc Pouly
arXiv Computer Vision
Aug 24

When does fusing hand-crafted knowledge with learned representations pay? A cost-normalized benchmark of stacking, substitution, and interference

arXiv:2608.21098v1 Announce Type: new Abstract: Fusing prior knowledge with data-driven learning is attractive where data is scarce, yet no controlled account says when it helps, is redundant, or har...

By Ahmad AlMughrabi, Albert Clop, Benjamin Busam, Ricardo Marques, Petia Radeva
arXiv Machine Learning
Aug 26

The Blending Ratio Is Not Where the Performance Is: Diagnosing Prototype Blending for Few-Shot Adaptation of Vision-Language Models

The paper investigates the blending ratio used in few‑shot adaptation of vision‑language models, which combines a zero‑shot text prototype with the mean of labeled image features. It shows that the theoretically optimal ratio—derived from a closed‑form mean‑squared error minimizer—does not align with the ratio that actually maximizes performance, falling short by an average of 8.5 points. Moreover, a leave‑one‑out estimate on the support set achieves near‑oracle performance, and validation‑free linear probes outperform even oracle‑tuned blends, indicating that the hyperparameter can be set near‑optimally without external validation data.

By Liangzhi Li, Bowen Wang, Yiming Qian, Thorsten Neumann, Xia Xie, Guangshun Li
arXiv Computer Vision
Sep 14

Beyond Argmax: A Mechanistic Study of Semantic Retention in Frozen Foundation-Model Composition for Generalized Few-Shot 3D Segmentation

The paper investigates how much semantic information is lost when frozen foundation models are combined for few‑shot 3D segmentation. By varying the number of retained semantic alternatives before fusion, the authors show that keeping the full distribution of class scores yields higher harmonic‑mean IoU than collapsing to a single class. Experiments on ScanNet200 and ScanNet++ confirm that full‑distribution fusion consistently outperforms top‑1 and other operators, and that most useful information is recovered by retaining a compact set of plausible alternatives.

By Silas Kwabla Gah, Ebenezer Owusu