arXiv AI By Issar Tzachor, Michael Green, Rami Ben-Ari

VidMsg: A Benchmark for Implicit Message Inference in Short Videos

Read the original on arXiv AI →

arXiv:2606. 03635v1 Announce Type: cross Abstract: Understanding short online videos involves more than identifying visible objects and actions; video makers often include an underlying message or purpose in the clip.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Jun 19

VCG: A Multimodal Retrieval Framework for E-Commerce Video Feeds under Extreme Cold-Start Conditions

arXiv:2606. 19627v1 Announce Type: cross Abstract: The digital commerce landscape is shifting from static, search-driven catalogs to dynamic, immersive video feeds.

By Katya Mirylenka, Egor Malykh, Mahdyar Ravanbakhsh, Michael Gygli, Marco-Andrea Buchmann, Andrew Dzhoha, Svitlana Borzenko, Francesca Catino, Mohamed Gaafar, Maarten Versteegh, Thomas Kober, Dario d'Andrea, Ellie Langhans
arXiv Computer Vision
Sep 24

MultiVENT-Raw: A Benchmark for Retrieval and Reasoning over Raw Videos

MultiVENT‑Raw is a new multilingual benchmark comprising nearly 120,000 raw videos—continuous footage from cell phones, hand‑held cameras, or CCTV—totaling over 5,300 hours. The dataset includes 130 events and 222 event‑centric queries, along with human‑annotated relevance judgments and extracted key facts for relevant videos. It supports two tasks: retrieving videos relevant to a query event and generating a coherent report summarizing event‑related videos for a target user, with baseline models showing these tasks remain challenging.

By Reno Kriz, David Etter, Alexander Martin, Cameron Carpenter, Debashish Chakraborty, Hannah Recknor, Reihaneh Iranmanesh, Matthew Maciejewski, Kenton Murray, Eugene Yang, Benjamin Van Durme, Aaron Steven White, Andrew Yates, William Walden
arXiv Computer Vision
Aug 27

AdaVDR: Adaptive Tool Use and Reflection for Video Deep Research

AdaVDR is an adaptive video deep research agent that selects and reflects on tool usage based on the task and the model’s capabilities. It constructs a specialized data pipeline to generate high‑quality QA pairs and uses model‑conditioned filtering to remove unnecessary tool calls. The agent is trained with supervised fine‑tuning and reinforcement learning, achieving top performance on the VDR‑EE benchmark and significant gains on VideoDR.

By Xintong Zhang, Xiaomeng Fan, Shilin Yan, Ekko He, Zicheng Liu, Zijian Zou, Guannan Zhang, Yuwei Wu, Zhi Gao, Hongwei Xue
arXiv AI
Sep 21

VidOmni-Bench: A Benchmark for Fine-Grained Video Understanding via Spatio-Temporal Event Verification across Complexity and Duration

VidOmni-Bench is a new benchmark for fine‑grained video understanding that asks models to verify whether each event in dense video captions is supported by the video. It contains 500 videos covering five complexity types and durations from 4 seconds to 90 minutes, and uses human‑verified sentence‑level labels to create hard negatives. Experiments show that Video‑LLMs often hallucinate events, struggle to detect incorrect descriptions, and exhibit varying weaknesses depending on video complexity and duration.

By Changbeen Kim, Junwon Chang, Kipyo Kim, Risa Shinoda, Kuniaki Saito, Donghyun Kim