arXiv:2609.00591v1 Announce Type: new
Abstract: An image may be worth a thousand words, but most captioning models describe it in only a few. Modern vision-language models produce fluent high-level c...
By Suryaansh Jain, Rahasya Barkur, Vishal G, Ryan Rossi, Franck Dernoncourt, Jack Wang, Koustava Goswami, Nedim Lipka, Puneet Mathur, Samyadeep Basu, Seunghyun Yoon
arXiv:2609.36893v1 Announce Type: new
Abstract: Detailed image captioning requires accurate and comprehensive descriptions of fine-grained visual content, yet caption quality spans factual accuracy,...
By Zhenwen Ji, Lei Jin, Shanyong Wang, Jiaming Lu, Chengqiang Lu, Yi Wu, Yao Hu, Lizhen Cui, Yanyu Xu
arXiv:2606. 00910v1 Announce Type: cross Abstract: Composed Video Retrieval (CoVR) seeks the target video that results from applying a free-form textual modification to a reference video.
By Ali Alavi
arXiv:2607. 03143v1 Announce Type: cross Abstract: Vision-language alignment powers open-vocabulary recognition, retrieval, and LVLM grounding, yet natural captions are often underspecified, making similarity brittle and overly confident under paraphrase and omitted details.
By Chengzhen Yu, Canran Xiao, Siyuan Ma, Yang Liu
The paper introduces Preserve-and-Compose Training (PACT) for composed image retrieval, a task where a query image is modified by a textual instruction while preserving visual content from a reference image. PACT learns from image–text–text triplets, using target captions for supervision and visual evidence from the source image to maintain relevant details, without requiring target images or gallery updates. The authors also propose Chord scoring, which blends target similarity with source-relative directional agreement in a frozen image space, and demonstrate that this combined approach yields strong retrieval performance across multiple zero-shot CIR benchmarks and various backbones.
By Sehyun Kwon
Object hallucination in multimodal large language models arises when language priors and corpus co-occurrence bias outweigh the visual evidence, with nothing tying an individual object mention to what the image shows. Most remedies intervene at decoding time without training, yet under a unified protocol their benefit is confined to short captions;supervised fine-tuning (SFT) on a detail- rich corpus lengthens captions, but over forty percent still name absent objects.