As a fundamental image processing task, object Re-Identification (ReID) aims to retrieve objects across non-overlapping cameras. Recently, with the development of deep learning, significant advancemen...
arXiv:2609.24539v1 Announce Type: new
Abstract: Multi-modal object Re-Identification (ReID) benefits from complementary information across heterogeneous imaging modalities. To further enrich semantic...
By Weixiang Zhou, Yuhao Wang, Xingguo Xu, Weizhen Zhou, Zhixun Su, Jinshan Pan, Cong Wang
arXiv:2606. 02242v1 Announce Type: cross Abstract: The joint optimization of image-based (I2I) and text-based (T2I) person re-identification (ReID) is hindered by modality discrepancies and conflicting training objectives, leading to suboptimal shared representations.
By Karina Kvanchiani, Timur Mamedov
The joint optimization of image-based (I2I) and text-based (T2I) person re-identification (ReID) is hindered by modality discrepancies and conflicting training objectives, leading to suboptimal shared representations. While I2I ReID focuses on identity-level invariance across images of the same person, T2I ReID is driven by instance-specific textual descriptions tied to unique visual traits.
arXiv:2609.14419v1 Announce Type: cross
Abstract: Person re-identification (ReID) is essential for multi-camera surveillance and tracking, yet remains difficult due to viewpoint and illumination chan...
By Leon Fernando, C Dombawala, P. Hettigoda, Vanodhya G. Warnasooriya, Ishara Neranjana, Rashmika Nawaratne
arXiv:2608. 18339v1 Announce Type: cross Abstract: Vision-language models (VLMs) have demonstrated remarkable zero-shot capabilities yet remain sensitive to real-world distribution shifts during inference.
By Qi Yu, Zhichen Zeng, Katherine Tieu, Xiyuan Yang, Ruizhong Qiu, Yuchen Yan, Lihui Liu, Yanjun Zhao, Lingjie Chen, Jingrui He, Hanghang Tong
UVU is a vision-language unified autoregressive framework that integrates visual supervision directly into the pre-training stage of multimodal large language models. By using continuous visual encoding and a large-scale iterative hierarchical clustering algorithm to build a pixel-level visual codebook, UVU enables lossless representation of visual inputs and autoregressive generation of pixel-level image tokens alongside textual tokens. This approach synergizes pixel-level visual perception with semantic-level visual understanding, allowing models to internalize visual reconstruction capabilities and improve multimodal understanding performance.
By Zhehan Kan, Xinghua Jiang, Yubo Zhu, Yanlin Liu, Xiaochen Yang, Zhixiang Wei, Shifeng Liu, Qingmin Liao, Wenming Yang, Xin Li, Yinsong Liu, Deqiang Jiang, Xing Sun
arXiv:2609.06967v1 Announce Type: cross
Abstract: Ensuring effective transfer learning for vision-language models without compromising their generalization performance is crucial. However, many exist...
By Seungmin Oh, Seunghun Kang, Jongbin Ryu
arXiv:2606. 16799v1 Announce Type: cross Abstract: Existing vision-language model (VLM)-based AI-generated image quality assessment (AIGIQA) methods suffer from a fundamental semantic-distortion dimensional conflict: monolithic representations optimized for semantic discrimination inherently entangle compositional understanding with low-level perceptual sensitivity, rendering them blind to fine-grained quality degradations.
By Zijie Meng
arXiv:2603.11521v2 Announce Type: replace-cross
Abstract: Unsupervised Camouflaged Object Detection (UCOD) remains a challenging task due to the high intrinsic similarity between target objects and t...
By Shuo Jiang, Gaojia Zhang, Min Tan, Yufei Yin, Gang Pan
arXiv:2409. 10094v3 Announce Type: replace-cross Abstract: Out-of-Distribution (OoD) detection aims to justify whether a given sample is from the training distribution of the classifier-under-protection, i.
By Kun Fang, Zuopeng Yang, Haibo Hu, Xiaolin Huang, Jie Yang, Qinghua Tao
arXiv:2606. 09362v1 Announce Type: cross Abstract: Re-Identification (ReID) in autonomous driving is typically formulated as a visual matching problem, where observations of vehicles, pedestrians, and cyclists are associated across time, frames, or camera views using learned appearance embeddings, often complemented by motion, geometric, or multimodal cues.
By Eduardo Borges, Manuel Abreu, Lu\'is Garrote, Urbano J. Nunes