arXiv:2607. 00620v1 Announce Type: cross Abstract: Generalized Category Discovery (GCD) aims to recognize known classes while autonomously discovering novel ones in open-world settings.
By Boyang Dai, Chaoqi Chen, Yizhou Yu
The paper introduces Auto-Comp, a fully automated, concept-driven pipeline that generates photorealistic compositional benchmarks for vision‑language models. Auto‑Comp creates paired Minimal and Contextual samples for each concept, enabling isolation of core binding abilities from visio‑linguistic complexity. Evaluations across 25 models reveal consistent failures in attribute and relational binding, with context helping relational tasks but hindering attribute tasks due to visual clutter.
By Cristian Sbrolli, Toshihiko Yamasaki, Matteo Matteucci
The paper introduces G2D, a training‑free framework that combines a discriminative model (CLIP) for broad candidate retrieval with a generative vision‑language model for fine‑grained, image‑grounded verification. By using CLIP’s top‑K shortlist and a structured prior from candidate names and probabilities, G2D focuses generative reasoning on uncertain samples, achieving an average accuracy of 68.85% across eight benchmarks—higher than both CLIP alone (59.35%) and the standalone generative model (63.11%). The approach also adapts to various generator configurations and extends to other models such as DCLIP, WaffleCLIP, and CuPL.
By Zehua Hao, Fang Liu, Qinliang Wang, Yaoyang Du, Xinyan Huang, Puhua Chen
arXiv:2508. 12745v2 Announce Type: replace-cross Abstract: Image set classification (ISC), which can be viewed as a task of comparing similarities between sets consisting of unordered heterogeneous images with variable quantities and qualities, has attracted growing research attention in recent years.
By Xizhan Gao, Wei Hu
HiPerViT is a compact vision-only architecture that injects an explicit second-order statistical prior into a transformer-based pipeline for texture recognition. It combines global and local image views with a compact bilinear descriptor encoded as a statistical token, and integrates this token with first-order spatial representations through Perceiver-style latent distillation. Across six texture recognition benchmarks, HiPerViT consistently outperforms strong vision-only baselines, achieving notable gains on DTD, GTOS-Mobile, and 1200Tex, and the improvements are largely independent of backbone depth or fusion topology.
By Jo\~ao Pedro C. A. de S\'a, Odemir Martinez Bruno
arXiv:2606. 24716v1 Announce Type: cross Abstract: Sparse autoencoders (SAEs) are increasingly used to extract interpretable concepts from vision and vision language models, yet existing evaluation methods largely rely on proxy metrics or qualitative inspection rather than measuring semantic correspondence.
By Jonas Klotz, Cassio F. Dantas, Pallavi Jain, Diego Marcos, Beg\"um Demir
arXiv:2607. 22919v1 Announce Type: cross Abstract: Multimodal embedding spaces in models like CLIP enable powerful capabilities such as semantic similarity retrieval and cross-modal zero-shot classification.
By Joseph Fioresi, Fabian Caba Heilbron, Pankaj Nathani, Mubarak Shah, Kushal Kafle
arXiv:2509. 09151v2 Announce Type: replace-cross Abstract: Research in video understanding has advanced rapidly, driven by increasingly diverse datasets and more powerful model architectures.
By Lei Wang, Syuan-Hao Li, Piotr Koniusz, Yongsheng Gao
arXiv:2609.21522v1 Announce Type: new
Abstract: Recent pre-trained foundation models provide rich multi-modal priors for downstream 3D vision tasks. However, the effectiveness of these representation...
By Hang Cheng, Yan Chen, Mingyu Fan, Long Zeng
G2D is a training‑free framework that combines a discriminative model (CLIP) for broad candidate retrieval with a generative vision‑language model for fine‑grained, image‑grounded verification. By using CLIP’s top‑K shortlist and a structured prior from candidate names and probabilities, G2D focuses generative reasoning on uncertain samples, employing fixed confidence routing, entropy‑adaptive candidate sizing, and trie‑constrained decoding to produce a single valid output. Across eight benchmarks, G2D achieves an average accuracy of 68.85%, outperforming both CLIP (59.35%) and the standalone generative model (63.11%), and it also transfers effectively to other models such as DCLIP, WaffleCLIP, and CuPL.
arXiv:2606. 15134v1 Announce Type: cross Abstract: Vision encoders for retrieval are typically trained with class-label supervision: each training pair reduces to a scalar that uniformly pushes the embedding apart or pulls it together, as if every visual attribute either differed or matched.
By Shubhang Bhatnagar, Dheeraj Baiju, Narendra Ahuja
The paper introduces a composition‑aware pretraining framework for geospatial foundation models that explicitly encodes fractional land‑cover mixtures as histogram targets for each satellite image cell. By using Earth Mover’s Distance to distill these composition targets into a 36.8 M‑parameter backbone, the authors demonstrate significant improvements on region‑level tasks such as zero‑shot image retrieval and scene classification, while maintaining competitive performance on fine‑grained tasks like segmentation and object detection. The method outperforms larger models (SatMAE and Prithvi‑EO‑2.0) and achieves a 55.6 % relative boost on the ForestNet‑12 dataset, evidencing the benefit of explicit composition modeling.
By Aryan Kashyap Naveen, Abhishek Srinivas, Pranav Moothedath, Shrutilipi Bhattacharjee