The paper introduces a Dual-Transformer architecture with Cross-Attention for multi-camera view recommendation, achieving a 56.60% Precision@0.5 on the TVMCE dataset, surpassing the previous best of 37.16%. The model separates temporal encoding of past frames from candidate view querying, and an ablation study shows the SwinV2 backbone yields 69.65% Precision@0.5. Fine‑tuning with as little as 20% of a target video improves precision, suggesting efficient personalization for specific editing styles.
By Josep Cabacas-Maso, Carles Ventura, Ismael Benito-Altamirano
The paper introduces SGWIB, a single‑modal video highlight detection framework that applies an information‑bottleneck approach while preserving inter‑segment temporal structure through a new Sliced Gromov‑Monge Gap regularizer. It also proposes Home‑Away‑Related Contextual Pseudo‑Labels and a contextual disentanglement module to mitigate sports‑specific bias. Experiments on MrHiSum and MoSu datasets show SGWIB outperforms existing methods on multiple ranking and accuracy metrics.
By Hanjuan Huang, Yung-Chieh Yeh, Hsing-Kuo Pao
Multi-camera systems are foundational to modern media production, and multi-camera editing is a critical task. This involves the proper selection of the appropriate camera view at each moment. In this...
EviDETR is a new framework for joint video moment retrieval and highlight detection that preserves query‑relevant temporal evidence throughout its pipeline. It introduces three key components: Semantic‑aware Feature Reweighting (SFR) to enhance clip representations, a Temporal Top‑2 Mixture‑of‑Experts (TTop2MoE) decoder for query‑adaptive refinement, and MR‑to‑HD (MR2HD) fusion to transfer retrieval evidence to highlight prediction. Using CLIP+SlowFast features, EviDETR achieves state‑of‑the‑art performance on QVHighlights and demonstrates strong cross‑dataset transferability on TACoS and Charades‑STA.
By Haoran Sun, Yufan Li, Qichen Zhang, Haoran Zhao, Shuqi Wang
Video highlight detection aims to identify temporally important segments that capture the most informative or engaging events in a video. Reliable prediction therefore requires not only discriminative...
arXiv:2606. 05702v1 Announce Type: new Abstract: Recent advancements in Vision-Language Models (VLMs) have significantly enhanced their ability to interpret complex visual semantics, yet their capacity for chronological reasoning remains under-explored.
By Haoyu Zhou, Qing Qing, Caichong Li, Qixin Zhang, Yongcheng Jing, Ziqi Xu, Juncheng Hu, Xikun Zhang, Renqiang Luo
The paper introduces Intrinsic Temporal Adaptation (ITA) for Partially Relevant Video Retrieval (PRVR), a task that seeks untrimmed videos containing moments relevant to a text query. ITA employs a Backbone-Internal Temporal Adaptation that lets the final visual transformer layers attend to neighboring frames, creating temporally aware embeddings while keeping CLIP frozen. Additionally, an Affinity-Weighted Gradient Propagation technique softly aggregates top‑k frames based on text‑frame affinities to better handle the weakly supervised nature of PRVR, leading to state‑of‑the‑art performance and more accurate frame‑level evidence retrieval.
By Hyun Seok Seong, Woojin Jun, SuBeen Lee, Jae-Pil Heo
arXiv:2607. 13421v1 Announce Type: cross Abstract: Spatio-Temporal Video Grounding (STVG) aims to retrieve the visual trajectory of a specific object from a video stream as described by a natural language expression.
By Kai Chen, Ming Dai, Wenxuan Cheng, Wankou Yang
Recent advancements in Vision-Language Models (VLMs) have significantly enhanced their ability to interpret complex visual semantics, yet their capacity for chronological reasoning remains under-explored. In this paper, we introduce a novel benchmark specifically designed to evaluate how VLMs perceive and reason about chronological information within and across images.
arXiv:2603.14733v2 Announce Type: replace
Abstract: Multimodal Large Language Models have achieved strong performance in single-video understanding, yet their ability to reason across multiple videos...
By Yue Zhang, Liqiang Jing, Jia Li, Yapeng Tian, Xinya Du, Yunhui Guo, Vibhav Gogate
OpenSAL360 is an open‑source platform that enables scalable, low‑cost collection of 360° video saliency data using only a standard screen, mouse, and internet connection. It bypasses the need for VR headsets, allowing parallel data collection from crowdsourced assessors. The authors validated the protocol against seven VR eye‑tracking datasets, performed ablation studies, and released a new dataset of 500 omnidirectional videos annotated by over 2,000 assessors, the largest in the field to date.
By Alexey Bryncev, Andrey Moskalenko, Kira Shilovskaya, Ivan Kosmynin, Dmitriy Vatolin
arXiv:2607. 17994v1 Announce Type: cross Abstract: Video understanding has become more and more important with the growth of Artificial Intelligence (AI) for video generation.
By Rui Chu, Yingjie Lao