arXiv:2606. 02242v1 Announce Type: cross Abstract: The joint optimization of image-based (I2I) and text-based (T2I) person re-identification (ReID) is hindered by modality discrepancies and conflicting training objectives, leading to suboptimal shared representations.
By Karina Kvanchiani, Timur Mamedov
arXiv:2609.12965v1 Announce Type: new
Abstract: Text-based person search (TBPS) aims to retrieve images of a target person from a large image gallery based on a given natural language description. Mo...
By Mang Ye, Yucheng Ji, Yang Bai, Min Cao, Siyuan Chai, Bo Du, Min Zhang
arXiv:2609.14419v1 Announce Type: cross
Abstract: Person re-identification (ReID) is essential for multi-camera surveillance and tracking, yet remains difficult due to viewpoint and illumination chan...
By Leon Fernando, C Dombawala, P. Hettigoda, Vanodhya G. Warnasooriya, Ishara Neranjana, Rashmika Nawaratne
arXiv:2609.36894v1 Announce Type: new
Abstract: As a fundamental image processing task, object Re-Identification (ReID) aims to retrieve objects across non-overlapping cameras. Recently, with the dev...
By Yingquan Wang, Pingping Zhang, Dong Wang, Huchuan Lu
arXiv:2603.02767v4 Announce Type: replace-cross
Abstract: Image--text contrastive pretraining has become a dominant paradigm for visual representation learning, yet existing methods often yield repre...
By Hanpeng Liu, Zidan Wang, Shuoxi Zhang, Zonglin Zhao, Zihao Bo, Rinyoichi Takezoe, Kaiwen Long, Yaqian Li, Kun He
As a fundamental image processing task, object Re-Identification (ReID) aims to retrieve objects across non-overlapping cameras. Recently, with the development of deep learning, significant advancemen...
arXiv:2506. 08774v2 Announce Type: replace-cross Abstract: Different machine learning models can represent the same underlying concept in different ways.
By Fan Xu, Luis A. Leiva
Traditional multimodal representation learning and generation are two stages: a contrastive or self-supervised visual encoder is trained first, followed by a separate downstream generative model. This...
Subject-driven personalized text-to-image generation requires a pretrained diffusion model to acquire a specific subject from a few reference images while preserving subject identity, following novel text prompts, and maintaining sample diversity. Existing optimization-based methods instantiate subject adaptation through full fine-tuning, textual embedding optimization, or low-rank parameter updates; PaRa further constrains personalization from the perspective of parameter rank reduction.
The paper investigates how the modality gap— the separation between image and text representations in contrastive vision‑language models—affects different downstream tasks. By showing that a single dominant direction accounts for most of the image‑text mean separation, the authors explain why reducing or removing this gap can improve zero‑shot classification, degrade retrieval, or restore performance depending on the task. The study provides a geometric framework that clarifies when and why gap interventions should be applied in vision‑language systems.
By Aditya Sharma, Divya Saxena
FLAT (Flexible‑Length Aligned Transmodal representations) is a joint multimodal pre‑training framework that learns a shared encoder for images and text, producing 1‑D continuous embeddings that can be directly used by downstream generative decoders. By combining contrastive alignment with bidirectional cross‑modal generative objectives, FLAT yields representations that are both discriminative and generative, enabling cross‑modal retrieval and generation with a single pre‑training stage. The model achieves strong performance on T2I generation (GenEval 71.1), image captioning (BLEU‑4 40.5, CIDEr 138.6), and retrieval tasks (Recall@5 86.8/75.8 on MS‑COCO, 98.3/93.6 on Flickr30K), and supports linear interpolation, latent space arithmetic, and zero‑shot composed retrieval.
By Guangyu Sun, Shlok Kumar Mishra, Wentao Bao, Robert Zhenheng Yang, Xiao Wang, Xiyuan Wang, Yujunrong Ma, Chen Yuan, Max Xiangjun Fan, Jun Xiao, Jianpeng Cheng
The paper introduces Redemption Score (RS), a multi‑modal evaluation framework for image captioning that combines three complementary signals: Mutual Information Divergence for global image‑text alignment, DINO‑based perceptual similarity of cycle‑generated images for visual grounding, and LLM text embeddings for contextual similarity to human references. RS fuses these signals to provide a more holistic assessment, achieving a Kendall‑τ of 58.42 on Flickr8k and outperforming most prior methods. The framework demonstrates consistent performance across Conceptual Captions and MS COCO, offering a robust evaluation that captures both visual accuracy and text quality.
By Ashim Dahal, Ankit Ghimire, Saydul Akbar Murad, Nick Rahimi