arXiv Computer Vision

TRINITY: A Multi-Perspective Benchmark for Personal-Style Video Highlight Detection

arXiv Computer Vision
Aug 27

A Dual-Transformer for Multi-Camera View Recommendation

The paper introduces a Dual-Transformer architecture with Cross-Attention for multi-camera view recommendation, achieving a 56.60% Precision@0.5 on the TVMCE dataset, surpassing the previous best of 37.16%. The model separates temporal encoding of past frames from candidate view querying, and an ablation study shows the SwinV2 backbone yields 69.65% Precision@0.5. Fine‑tuning with as little as 20% of a target video improves precision, suggesting efficient personalization for specific editing styles.

By Josep Cabacas-Maso, Carles Ventura, Ismael Benito-Altamirano
arXiv AI
Sep 15

SGWIB:Sliced Gromov-Wasserstein Information Bottleneck for Video Highlight Detection

The paper introduces SGWIB, a single‑modal video highlight detection framework that applies an information‑bottleneck approach while preserving inter‑segment temporal structure through a new Sliced Gromov‑Monge Gap regularizer. It also proposes Home‑Away‑Related Contextual Pseudo‑Labels and a contextual disentanglement module to mitigate sports‑specific bias. Experiments on MrHiSum and MoSu datasets show SGWIB outperforms existing methods on multiple ranking and accuracy metrics.

By Hanjuan Huang, Yung-Chieh Yeh, Hsing-Kuo Pao
arXiv Computer Vision
6d ago

EviDETR: Preserving Query-Relevant Temporal Evidence for Moment Retrieval and Highlight Detection

EviDETR is a new framework for joint video moment retrieval and highlight detection that preserves query‑relevant temporal evidence throughout its pipeline. It introduces three key components: Semantic‑aware Feature Reweighting (SFR) to enhance clip representations, a Temporal Top‑2 Mixture‑of‑Experts (TTop2MoE) decoder for query‑adaptive refinement, and MR‑to‑HD (MR2HD) fusion to transfer retrieval evidence to highlight prediction. Using CLIP+SlowFast features, EviDETR achieves state‑of‑the‑art performance on QVHighlights and demonstrates strong cross‑dataset transferability on TACoS and Charades‑STA.

By Haoran Sun, Yufan Li, Qichen Zhang, Haoran Zhao, Shuqi Wang
arXiv AI
Jun 6

Seeing Time: Benchmarking Chronological Reasoning and Shortcut Biases in Vision-Language Models

arXiv:2606. 05702v1 Announce Type: new Abstract: Recent advancements in Vision-Language Models (VLMs) have significantly enhanced their ability to interpret complex visual semantics, yet their capacity for chronological reasoning remains under-explored.

By Haoyu Zhou, Qing Qing, Caichong Li, Qixin Zhang, Yongcheng Jing, Ziqi Xu, Juncheng Hu, Xikun Zhang, Renqiang Luo
arXiv Computer Vision
Sep 7

Intrinsic Temporal Adaptation of CLIP for Partially Relevant Video Retrieval

The paper introduces Intrinsic Temporal Adaptation (ITA) for Partially Relevant Video Retrieval (PRVR), a task that seeks untrimmed videos containing moments relevant to a text query. ITA employs a Backbone-Internal Temporal Adaptation that lets the final visual transformer layers attend to neighboring frames, creating temporally aware embeddings while keeping CLIP frozen. Additionally, an Affinity-Weighted Gradient Propagation technique softly aggregates top‑k frames based on text‑frame affinities to better handle the weakly supervised nature of PRVR, leading to state‑of‑the‑art performance and more accurate frame‑level evidence retrieval.

By Hyun Seok Seong, Woojin Jun, SuBeen Lee, Jae-Pil Heo
Hugging Face Trending Papers
Jun 4

Seeing Time: Benchmarking Chronological Reasoning and Shortcut Biases in Vision-Language Models

Recent advancements in Vision-Language Models (VLMs) have significantly enhanced their ability to interpret complex visual semantics, yet their capacity for chronological reasoning remains under-explored. In this paper, we introduce a novel benchmark specifically designed to evaluate how VLMs perceive and reason about chronological information within and across images.

arXiv Computer Vision
Sep 21

OpenSAL360: Open-Source Crowdsourcing Platform for Omnidirectional Video Saliency Collection

OpenSAL360 is an open‑source platform that enables scalable, low‑cost collection of 360° video saliency data using only a standard screen, mouse, and internet connection. It bypasses the need for VR headsets, allowing parallel data collection from crowdsourced assessors. The authors validated the protocol against seven VR eye‑tracking datasets, performed ablation studies, and released a new dataset of 500 omnidirectional videos annotated by over 2,000 assessors, the largest in the field to date.

By Alexey Bryncev, Andrey Moskalenko, Kira Shilovskaya, Ivan Kosmynin, Dmitriy Vatolin