arXiv:2607. 17994v1 Announce Type: cross Abstract: Video understanding has become more and more important with the growth of Artificial Intelligence (AI) for video generation.
By Rui Chu, Yingjie Lao
arXiv:2608.24845v1 Announce Type: cross
Abstract: We present LAION-BVD, a large-scale open video dataset for multimodal learning, which contains 1.3B platform-specific video URLs collected from Commo...
By Andreas Hochlehnert, Marianna Nezhurina, Mehdi Cherti, Andrej Radonjic, Thadd\"aus Wiedemer, Christoph Schuhmann, Romain Beaumont, Wieland Brendel, Bernhard Sch\"olkopf, A. Sophia Koepke, Jenia Jitsev, Matthias Bethge
arXiv:2610.01022v1 Announce Type: new
Abstract: Memory-attention-based Video Instance Segmentation (VIS) methods have demonstrated strong zero-shot tracking capability, yet their substantial memory r...
By Arash Rocky, Q. M. Jonathan Wu
arXiv:2505.01583v2 Announce Type: replace
Abstract: Understanding causal event relationships and achieving fine-grained temporal grounding in videos remain challenging for vision-language models (VLM...
By Jen-Hao Cheng, Yi-Hao Peng, Huapeng Zhou, Vivian Wang, Huayu Wang, Hsiang-Wei Huang, Wenhao Chai, Hou-I Liu, Kuang-Ming Chen, Cheng-Yen Yang, Yi-Ling Chen, Vibhav Vineet, Qin Cai, Jenq-Neng Hwang
TAME is a CLIP‑based framework for Text‑Video Retrieval that incorporates temporal modeling through three key innovations: sparse Mixture‑of‑Experts layers with frame‑consistent routing, Frame‑Temporal tokens that aggregate cross‑frame information, and a Cross‑Temporal Interaction and Aggregation module for refining frame‑wise similarities. These components enable the model to capture both local visual patterns and long‑range temporal dependencies, leading to consistent performance gains over CLIP‑based baselines on multiple TVR benchmarks, including a 4.0 R@1 improvement on MSR‑VTT. The code is publicly available on GitHub.
By Uicheol Jung, Juyoung Hong, Hojung Kwon, Yukyung Choi
We present LAION-BVD, a large-scale open video dataset for multimodal learning, which contains 1.3B platform-specific video URLs collected from CommonCrawl. From these, we download 80M videos with a t...
arXiv:2604. 10024v2 Announce Type: replace-cross Abstract: Long video summarization presents significant challenges for multimodal large language models (MLLMs), particularly in maintaining temporal fidelity over extended durations and producing summaries that are both semantically and temporally grounded.
By Alkesh Patel, Melis Ozyildirim, Ying-Chang Cheng, Ganesh Nagarajan
ShotFinder introduces a new benchmark for open‑domain video shot retrieval, formalizing editing requirements as keyframe‑oriented shot descriptions and adding five controllable constraints—temporal order, color, visual style, audio, and resolution. The benchmark comprises 1,210 high‑quality YouTube samples across 20 themes, generated with large models and verified by humans. A three‑stage retrieval pipeline—query expansion via video imagination, candidate video retrieval, and description‑guided shot localization—shows a notable performance gap to humans, especially for color and visual style constraints.
By Tao Yu, Haopeng Jin, Hao Wang, Shenghua Chai, Yujia Yang, Junhao Gong, Jiaming Guo, Minghui Zhang, Xinlong Chen, Zhenghao Zhang, Yuxuan Zhou, Yufei Xiong, Shanbin Zhang, Jiabing Yang, YiFan Zhang, Hongzhu Yi, Xinming Wang, Cheng Zhong, Xiao Ma, Zhang Zhang, Yan Huang, Liang Wang
arXiv:2509.22650v3 Announce Type: replace
Abstract: Most existing approaches to referring segmentation achieve strong performance only through fine-tuning or by composing multiple pre-trained models,...
By Anna Kukleva, Enis Simsar, Alessio Tonioni, Muhammad Ferjad Naeem, Federico Tombari, Jan Eric Lenssen, Bernt Schiele
VidOmni-Bench is a new benchmark for fine‑grained video understanding that asks models to verify whether each event in dense video captions is supported by the video. It contains 500 videos covering five complexity types and durations from 4 seconds to 90 minutes, and uses human‑verified sentence‑level labels to create hard negatives. Experiments show that Video‑LLMs often hallucinate events, struggle to detect incorrect descriptions, and exhibit varying weaknesses depending on video complexity and duration.
By Changbeen Kim, Junwon Chang, Kipyo Kim, Risa Shinoda, Kuniaki Saito, Donghyun Kim
arXiv:2510. 14904v4 Announce Type: replace-cross Abstract: Dense Video Object Captioning (DVOC) is the task of jointly detecting, tracking, and captioning object trajectories in a video, requiring the ability to understand spatio-temporal details and describe them in natural language.
By Gabriel Fiastre, Antoine Yang, Cordelia Schmid
The paper introduces Seeing Before Synthesizing (SBS), a weakly-supervised dense video captioning framework that uses a vision‑language model to generate frame‑level narratives for gaps between events and detect transitions based on semantic changes. SBS refines temporal masks by aligning transition points with vision‑language cues, rather than relying on rigidly placed synthetic captions. Experiments on ActivityNet Captions and YouCook2 show that SBS achieves state‑of‑the‑art results in both captioning and localization tasks.
By Ye-Chan Kim, Seunghee Choi, SeungJu Cha, Si-Woo Kim, Hwiseon Kim, Hyungee Kim, Dong-Jin Kim