arXiv:2609.36560v1 Announce Type: cross
Abstract: Object re-identification (ReID) faces a recurring challenge: different identities can share highly similar global appearances, while the cues that di...
By Zhiqi Li, Xiaowei Zhou, Zeyuan Sun, Feng Gao, Junyu Dong
arXiv:2607. 17157v1 Announce Type: cross Abstract: Multi-object tracking (MOT) aims to localize multiple objects in videos while preserving their identities over time.
By Yanrong Qin, Xiaoyan Cao, Yao Yao
arXiv:2607. 22068v1 Announce Type: cross Abstract: Multi-branch architectures and CNN-Transformer fusion have long been regarded as effective ways to improve vehicle re-identification (Re-ID) by combining complementary representations.
By Yu Wang, Hongyu Yang
arXiv:2609.14419v1 Announce Type: cross
Abstract: Person re-identification (ReID) is essential for multi-camera surveillance and tracking, yet remains difficult due to viewpoint and illumination chan...
By Leon Fernando, C Dombawala, P. Hettigoda, Vanodhya G. Warnasooriya, Ishara Neranjana, Rashmika Nawaratne
arXiv:2603.04163v2 Announce Type: replace
Abstract: Wildlife re-identification aims to recognise individual animals by matching query images to a database of previously identified individuals, based...
By Thanos Polychronou, Luk\'a\v{s} Adam, Viktor Penchev, Kostas Papafitsoros
arXiv:2607.09086v2 Announce Type: replace
Abstract: We present Subtoken Vision Transformer (SubViT), a selective image tokenization method for fine-grained visual recognition. Standard Vision Transfo...
By Jie Zhu, Ivy Zhang, Minchul Kim, Xiaoming Liu
MINER is a training‑free inference framework that enhances frozen dual‑encoder models for text‑to‑image retrieval when queries refer to small, visually subordinate objects in cluttered scenes. It augments the global image embedding with a bank of region‑level embeddings and applies hubness‑correcting similarity rescoring, thereby recovering visual evidence that global pooling underweights. The authors introduce ROCS, a benchmark derived from Flickr30K and MS COCO, and demonstrate that MINER improves retrieval performance across CLIP, SigLIP, and SigLIP 2 backbones on both ROCS and standard splits, attributing gains mainly to broader spatial coverage rather than precise crop placement.
By Abdulmalik Alquwayfili, Faisal AlMeshal, Jumanah Almajnouni, Huda Abdulhadi Alamri, Muhammad Kamran J Khan
arXiv:2609.09705v1 Announce Type: new
Abstract: Generalizable animal Re-Identification (ReID) aims to recognize individual animals across species with diverse morphologies and ecological contexts. Un...
By Shuoyi Chen, Yuejia Li, Mang Ye
arXiv:2607.02486v2 Announce Type: replace
Abstract: Descriptor-free visual localization eliminates high-dimensional descriptor storage, preserves scene privacy, and simplifies map maintenance, yet it...
By Yejun Zhang, Xinjue Wang, Zihan Wang, Esa Rahtu, Juho Kannala
Multi-vector vision-language retrieval preserves fine-grained visual evidence through maximum-similarity late interaction, but dense image-side tokens make storage and scoring expensive. Existing token compression methods reduce this cost, yet they can remove or collapse object- and region-level evidence that future query tokens may need to select.
VastMAT is a large‑scale multi‑animal tracking benchmark featuring 2,947 videos, 337 animal categories, and over 3.6 million bounding boxes with 22,883 identity trajectories. It emphasizes high‑quality, expert‑reviewed annotations and introduces Seen‑category and Unseen‑category evaluation protocols, revealing significant challenges in tracking unseen animals. The authors also propose a lightweight Center‑Distance‑Augmented Association module that boosts HOTA scores for existing MOT methods without extra training.
The paper introduces a one-stage end-to-end model for wildlife instance-level recognition that integrates detection and re-identification within a single latent space. It leverages DINOv2 for spatial geometry and MegaDescriptor for re-identification, while enhancing latent queries with prompt re-identification features. Preliminary results show a competitive mean average precision of 30.584% compared to the state-of-the-art two-stage approach of 44.89%, with qualitative evidence of effective bounding and identification of animal identities.
By Mufhumudzi Muthivhi, Jiahao Huo, Terence van Zyl, Fredrik Gustafsson