The paper presents a new approach to automatic audio description (AD) that treats the task as a constrained global optimization problem. It jointly decides what visual content is narratively important, when it can be spoken without overlapping dialogue, and how to phrase it within time limits. Using large language models for salience estimation and a mixed‑integer linear program for scheduling, the system outperforms prior methods on the REFRAMED benchmark, especially in temporal placement and narrative relevance.
By Igor Sterner, Mirella Lapata, Alex Lascarides, Frank Keller
arXiv:2607. 11798v1 Announce Type: cross Abstract: Long-form audio description (AD) requires more than describing visible actions: it must preserve characters, events, relationships, and story context across scenes so that blind and low-vision (BLV) audiences can follow a film.
By Seung Hyun Hahm, Minh T. Dinh, SouYoung Jin
arXiv:2609.36218v1 Announce Type: cross
Abstract: Large language models are increasingly evaluated in specialized domains such as law, medicine, software engineering, and cybersecurity, yet film rema...
By Mir Tafseer Nayeem, Susmoy Chakraborty, Davood Rafiei
arXiv:2608.21430v1 Announce Type: new
Abstract: Multimodal language models increasingly show promise for enabling the large-scale computational analysis of film, opening up new avenues for learning a...
By David Bamman, Kent K. Chang, Allison Cooper, Juishan Hsu, Reina Kushihashi, Madison Mar, Arnav Podichetty, Rachael Samberg, Ipek Nil Sancak, Yuhan Shao
Audio Description (AD) provides spoken narration of visual events during dialogue gaps, making movies accessible to visually impaired audiences. The problem requires determining both what (which visua...
The paper introduces Cue2Narrate, a two‑stage pipeline that jointly predicts what visual events to narrate and when to insert the narration in long, untrimmed movie clips. It uses a dual‑head audio‑visual localizer to identify visual cue and narration windows, followed by a LoRA‑adapted vision‑language model that generates concise audio descriptions, trained with a Description Ranking Loss. The authors also present the LongLSMDC benchmark, comprising up to 8‑minute clips, and show that Cue2Narrate outperforms video‑only and audio‑only baselines by 5–12 points in average mAP and improves AD generation over fine‑tuned base VLMs.
By Akshita Gupta, Aditya Arora, Federico Tombari, Marcus Rohrbach, Anna Rohrbach
arXiv:2606. 09846v1 Announce Type: cross Abstract: Visual art remains largely inaccessible to blind and low-vision (BLV) audiences due to brief or absent alt-text, which rarely conveys the sensory, spatial, or emotional qualities of an artwork.
By Vignesh Nagarajan
SonicCaps is a large-scale audio captioning dataset featuring approximately 15 million captions paired with 700,000 audio clips, created using the Qwen3-Omni multimodal language model. The dataset emphasizes diversity by generating around 24 captions per clip through structured prompt engineering and few-shot generation, covering main descriptions, rephrased variants, and semantic tags. Human evaluations rate SonicCaps higher than existing datasets, and training CLAP models on it improves audio retrieval and zero-shot classification across public and commercial benchmarks.
By Zineb Lahrichi, Marc Ferras, Ga\"el Richard, Geoffroy Peeters
arXiv:2608.30068v1 Announce Type: new
Abstract: Films communicate through deliberate creative choices, including lighting, color, composition, editing, dialogue, music, and sound. Humans naturally in...
By Kaishuu Shinozaki-Conefrey, Olivier Pascaud, Robin Courant, Xi Wang, Dimitris Samaras, Vicky Kalogeiton
AVSD-Scenes is a new dataset of 12,291 audio‑visual scene descriptions for urban environments, built from the TAU Urban Audio‑Visual Scenes dataset. The descriptions are generated by first creating modality‑specific text with Qwen2‑Audio‑7B and Qwen2.5‑VL‑7B, then merging them with large language models (Qwen3‑14B, Mistral‑Small‑3.2‑24B‑Instruct‑2506, Gemma‑3‑27B‑it) to produce multimodal narratives that combine auditory and visual cues. Benchmarks show that these multimodal descriptions improve semantic alignment, cross‑modal retrieval, and scene classification accuracy (up to 95.4%) while remaining discriminative even without explicit scene labels.
By Dhanunjaya Varma Devalraju, Arshdeep Singh, Mark D. Plumbley
The paper introduces MObyGaze, a dataset of 20 films annotated by experts for multimodal objectification, covering 6072 segments across 43 hours of video. It defines objectification through a structured thesaurus of 5 sub‑constructs and 11 concepts spanning visual, speech, and audio modalities. The authors formulate learning tasks, explore label diversity strategies, and benchmark vision, text, and audio models to demonstrate the task’s feasibility.
By Julie Tores, Elisa Ancarani, Lucile Sassatelli, Hui-Yin Wu, Clement Bergman, Lea Andolfi, Victor Ecrement, Remy Sun, Frederic Precioso, Thierry Devars, Magali Guaresi, Virginie Julliard, Sarah Lecossais
arXiv:2607. 02504v1 Announce Type: cross Abstract: Long-form TV dramas present a formidable challenge for comprehensive video understanding, where deciphering complex storyline often relies on \textbf{speaker recognition}, the task of accurately attributing each spoken utterance to its respective character.
By Yuxuan Li, Lingxi Xie, Xinyue Huo, Jihao Qiu, Jiacheng Shao, Pengfei Chen, Jiannan Ge, Kaiwen Duan, Qi Tian