The paper introduces CoVR‑R, a reason‑aware composed video retrieval system that, given a reference video and an edit instruction, retrieves a target video that satisfies the edit. It employs a zero‑shot reason‑then‑retrieve pipeline using Qwen3.5‑27B to generate structured descriptions and dense embeddings for gallery videos, and performs edit reasoning on the query to produce a target‑video description used as the query embedding. The method combines dense retrieval with a TF‑IDF branch over generated texts, fusing the rankings with split‑specific weights, achieving state‑of‑the‑art retrieval metrics on both validation and blind test splits.
By Dongqing Liu, Mengshi Qi, Hongwei Ji
arXiv:2609.10008v1 Announce Type: new
Abstract: Composed video retrieval (CoVR) searches a gallery for the target video that realizes a natural-language modification of a source clip. However, at gal...
By Dmitry Demidov, Muhammad Zaigham Zaheer, Omkar Thawakar, Abdelrahman Mohamed Shaker, Rao Anwer
SCOUT is a frozen‑encoder approach for sim‑to‑real text‑based person retrieval that predicts cross‑modal embeddings instead of fine‑tuning cross‑encoders. It uses a trainable predictor to map patch tokens from a frozen video encoder (V‑JEPA) into the embedding space of a frozen text encoder (EmbeddingGemma), guided by a bidirectional InfoNCE objective. The method achieves state‑of‑the‑art results on the AI City Challenge 2026 Track 4, with a full retrieve‑fuse‑rerank pipeline reaching 84.25 mAP@10 and a single frozen model alone scoring 60.63, while training costs are modest (≈95 GPU‑hours).
By Abdarahmane Traor\'e, Andy Couturier, \'Eric Hervet
arXiv:2606. 24477v1 Announce Type: cross Abstract: Video large language models (LLMs) are often constrained by computation and memory budgets, leading them to use reduced frame rates and spatial resolutions, which may cause them to miss critical information for question answering (QA).
By Yixuan Li, Guangzhi Sun, Yudong Yang, Wei Li, Zejun MA, Chao Zhang
MARS introduces a multi‑layer, multi‑slot embedding framework for text‑video retrieval that constructs adaptive representation slots by combining hidden states from different decoder layers. By comparing corresponding text and video slots and aggregating their similarities, MARS captures fine‑grained cues that single‑token embeddings miss. A hard‑negative‑aware slot specialization objective further encourages slots to focus on discriminative matching cues, leading to state‑of‑the‑art results on four benchmarks.
By Uicheol Jung, Juyoung Hong, Geuntaek Lim, Yukyung Choi
arXiv:2606. 13141v1 Announce Type: new Abstract: Retrieval-augmented generation is moving beyond text into long, egocentric video, where systems must select query-relevant chunks across multiple modalities and temporal granularities.
By Yuho Lee, Jisu Shin, Nicole Hee-Yeon Kim, Jihwan Bang, Juntae Lee, Kyuwoong Hwang, Fatih Porikli, Hwanjun Song
arXiv:2608.05707v2 Announce Type: replace
Abstract: Frame selection is essential for applying Large Multimodal Models (LMMs) to long videos due to severe frame redundancy and limited context windows....
By Wang Chen, Yu Chen, Xiang Wang, Shuai Li, Jinfa Huang, Xiawu Zheng
arXiv:2609.37426v1 Announce Type: cross
Abstract: Modern vision-language models (VLMs) have shown promising results in long-video understanding due to the rich semantic information they can capture....
By Arka Mukherjee, Kaleen Shrestha, Larissa Zhu, Maja Matari\'c
Edit-conditioned 3D scene retrieval pairs a reference 3D room with a natural-language modification and retrieves rooms from a corpus that satisfy the edit. Three lines of prior work each fall short on this task.
arXiv:2609.37345v1 Announce Type: new
Abstract: Streaming video understanding requires Video Large Language Models (Video-LLMs) to reason over continuous visual streams under causal constraints. As t...
By Xiang Hu, Jiazuo Yu, Lu Zhang, Yunzhi Zhuge, Huchuan Lu
arXiv:2607. 28986v1 Announce Type: cross Abstract: Zero-shot image captioning (ZIC) describes images without paired image-caption supervision during captioner training, relying on text-only corpora and frozen pretrained image-text scorers.
By Duy Tran Thanh, Thien-Phuc Doan, Long Nguyen-Vu, Ngo Tan Vu Khanh
arXiv:2606. 01485v1 Announce Type: cross Abstract: We describe our submission to the VRR Challenge @ CVPR 2026, built on the \emph{ImplicitQA} / \emph{VRR-QA} benchmark~\cite{implicitqa}: multiple-choice video question answering in which answers are deliberately \emph{not} observable in any single frame and must be inferred from spatial layout, motion, depth, viewpoint, causality, and social context across discontinuous frames of creative video.
By Ali Alavi