ShotFinder introduces a new benchmark for open‑domain video shot retrieval, formalizing editing requirements as keyframe‑oriented shot descriptions and adding five controllable constraints—temporal order, color, visual style, audio, and resolution. The benchmark comprises 1,210 high‑quality YouTube samples across 20 themes, generated with large models and verified by humans. A three‑stage retrieval pipeline—query expansion via video imagination, candidate video retrieval, and description‑guided shot localization—shows a notable performance gap to humans, especially for color and visual style constraints.
By Tao Yu, Haopeng Jin, Hao Wang, Shenghua Chai, Yujia Yang, Junhao Gong, Jiaming Guo, Minghui Zhang, Xinlong Chen, Zhenghao Zhang, Yuxuan Zhou, Yufei Xiong, Shanbin Zhang, Jiabing Yang, YiFan Zhang, Hongzhu Yi, Xinming Wang, Cheng Zhong, Xiao Ma, Zhang Zhang, Yan Huang, Liang Wang
arXiv:2609.10008v1 Announce Type: new
Abstract: Composed video retrieval (CoVR) searches a gallery for the target video that realizes a natural-language modification of a source clip. However, at gal...
By Dmitry Demidov, Muhammad Zaigham Zaheer, Omkar Thawakar, Abdelrahman Mohamed Shaker, Rao Anwer
MultiVENT‑Raw is a new multilingual benchmark comprising nearly 120,000 raw videos—continuous footage from cell phones, hand‑held cameras, or CCTV—totaling over 5,300 hours. The dataset includes 130 events and 222 event‑centric queries, along with human‑annotated relevance judgments and extracted key facts for relevant videos. It supports two tasks: retrieving videos relevant to a query event and generating a coherent report summarizing event‑related videos for a target user, with baseline models showing these tasks remain challenging.
By Reno Kriz, David Etter, Alexander Martin, Cameron Carpenter, Debashish Chakraborty, Hannah Recknor, Reihaneh Iranmanesh, Matthew Maciejewski, Kenton Murray, Eugene Yang, Benjamin Van Durme, Aaron Steven White, Andrew Yates, William Walden
arXiv:2606. 03635v1 Announce Type: cross Abstract: Understanding short online videos involves more than identifying visible objects and actions; video makers often include an underlying message or purpose in the clip.
By Issar Tzachor, Michael Green, Rami Ben-Ari
arXiv:2607. 00446v1 Announce Type: cross Abstract: As video corpora continue to expand in both scale and task complexity, there is increasing demand for approaches that retrieve relevant videos from large-scale corpora (inter-video reasoning) and subsequently perform fine-grained, query-conditioned tasks (intra-video reasoning) within the retrieved content, such as temporal grounding.
By Seohyun Lee, Seoung Choi, Dohwan Ko, Jongha Kim, Hyunwoo J. Kim
M3TR is a temporal retrieval‑enhanced multi‑modal framework for predicting micro‑video popularity. It introduces a Mamba‑Hawkes Process module to model user feedback as self‑exciting events, capturing long‑range temporal dependencies. A temporal‑aware retrieval engine then identifies historically relevant videos by combining multi‑modal content similarity with popularity trajectory similarity, augmenting the target video’s features for improved prediction accuracy.
By Jiacheng Lu, Weijian Wang, Mingyuan Xiao, Yang Hua, Tao Song, Bo Peng, Cheng Hua, Haibing Guan