arXiv Computer Vision

TAKE 85: Testing Audiovisual filmmaKer's intEnt across 85 Hours of Film

arXiv AI
Jun 30

MuseBench: Benchmarking Intent-Level Audiovisual Arts Understanding in MLLMs

arXiv:2606. 30026v1 Announce Type: cross Abstract: Audiovisual arts encompass diverse creative disciplines, including cinema, visual arts, stage performance, and game design, where artistic meaning arises from deliberate combinations of visual, auditory, and narrative elements (e.

By Yuxuan Fan, Gyusik Seo, Jing Hao, Jaemin Cho, Mohit Bansal, Jaehong Yoon
arXiv AI
Aug 3

COSI-Lab: Conference Living Lab for Modeling Multi-Perspective Multimodal Social Intention

arXiv:2607. 28649v1 Announce Type: cross Abstract: COSI-Lab presents a multimodal, multi-sensor dataset of an interdisciplinary scientific workshop containing 32 academics at an international conference.

By Zonghuan Li, Litian Li, Arthur Mercier, Gara Dorta, Balint Dioszegi, Jose Morales-Vargas, Chenxu Hao, Ivan Kondyurin, Vanessa Begemann, Nale Lehmann-Willenbrock, Bernd Dudzik, Saunaq Chakrabarty, Sotiris Vacanas, Laura Cabrera-Quir\'os, Anne L. J. ter Wal, Vitaliy Popov, Jorge Castro-God\'inez, Chirag Raman, Stephanie Tan, Hayley Hung
arXiv Computer Vision
Aug 27

MObyGaze: a film dataset of multimodal objectification densely annotated by experts

The paper introduces MObyGaze, a dataset of 20 films annotated by experts for multimodal objectification, covering 6072 segments across 43 hours of video. It defines objectification through a structured thesaurus of 5 sub‑constructs and 11 concepts spanning visual, speech, and audio modalities. The authors formulate learning tasks, explore label diversity strategies, and benchmark vision, text, and audio models to demonstrate the task’s feasibility.

By Julie Tores, Elisa Ancarani, Lucile Sassatelli, Hui-Yin Wu, Clement Bergman, Lea Andolfi, Victor Ecrement, Remy Sun, Frederic Precioso, Thierry Devars, Magali Guaresi, Virginie Julliard, Sarah Lecossais
arXiv AI
Aug 26

EXAM$^2$: $\underline{Ex}tending$ $\underline{A}udio$ $Understanding$ $in$ $\underline{M}ultilingual$ $and$ $\underline{M}ultimodal$ $Analysis$

EXAM$^2$ is a new benchmark for audio understanding that covers six languages and multiple modalities—speech, sound, music, mixed-audio, and visual images—providing 5,667 multiple-choice questions, 22,614 image instances, and 135,684 multilingual translations. It evaluates large audio language models (LALMs) and multimodal large language models (LLMs), revealing significant gaps in multilingual and cross‑modal performance. The authors also introduce Gemma3n-EXAM$^2$, a lightweight fusion model that improves multilingual results by up to 12.4% and multimodal results by 21.7% over a strong baseline.

By Jiawen Wang, Xiaoxue Gao, Zi Haur Pang, Nancy F. Chen