K9-Bench: Evaluating Multimodal LLMs on Canine-Centric Videos
arXiv:2607. 02680v1 Announce Type: cross Abstract: MLLMs have shown strong zero-shot capabilities across diverse inputs such as across images, video, audio, and text.
arXiv:2607. 06875v1 Announce Type: cross Abstract: Understanding and forecasting audience reactions to video content are crucial for improving content creation, recommendation systems, and media analysis.
arXiv:2607. 02680v1 Announce Type: cross Abstract: MLLMs have shown strong zero-shot capabilities across diverse inputs such as across images, video, audio, and text.
arXiv:2606. 19627v1 Announce Type: cross Abstract: The digital commerce landscape is shifting from static, search-driven catalogs to dynamic, immersive video feeds.
arXiv:2606. 15694v1 Announce Type: cross Abstract: Multimodal large language models (MLLMs) have demonstrated remarkable capabilities in understanding complex multimodal content.
arXiv:2608. 14391v1 Announce Type: cross Abstract: Recent video generators can fabricate realistic depictions of wars, disasters, public emergencies, and other real-world crises, creating substantial risks of misinformation.
arXiv:2606. 07433v1 Announce Type: cross Abstract: Video understanding is being rapidly transformed by multimodal large language models (MLLMs), as research moves from short clips to long, multimodal, and knowledge-intensive video scenarios.
arXiv:2603. 26772v2 Announce Type: replace-cross Abstract: Automated semantic annotation of broadcast television content presents distinctive challenges, combining structured audiovisual composition, domain-specific editorial patterns, and strict operational constraints.
arXiv:2607. 12774v1 Announce Type: cross Abstract: This article presents our results for the 11th Affective Behavior Analysis in-the-Wild (ABAW) competition.
arXiv:2603. 09731v3 Announce Type: replace-cross Abstract: Multimodal large language models (MLLMs) are increasingly considered as a foundation for embodied agents, yet it remains unclear whether they can reliably reason about the long-term physical consequences of actions from an egocentric viewpoint.
arXiv:2608. 13990v1 Announce Type: new Abstract: Driven by the attention economy, short-video Recommender Systems (RSs) are primarily optimized to maximize user engagement by promoting videos that capture attention within seconds.
arXiv:2509. 09151v2 Announce Type: replace-cross Abstract: Research in video understanding has advanced rapidly, driven by increasingly diverse datasets and more powerful model architectures.
arXiv:2601. 08828v2 Announce Type: replace-cross Abstract: Despite the rapid progress of video generation models, the role of data in influencing motion is poorly understood.
arXiv:2606. 14958v1 Announce Type: cross Abstract: We introduce the Massive Video Embedding Benchmark (MVEB), a 23-task benchmark for video embeddings spanning classification, zero-shot classification, clustering, pair classification, retrieval, and video-centric question answering.