arXiv:2606. 30026v1 Announce Type: cross Abstract: Audiovisual arts encompass diverse creative disciplines, including cinema, visual arts, stage performance, and game design, where artistic meaning arises from deliberate combinations of visual, auditory, and narrative elements (e.
By Yuxuan Fan, Gyusik Seo, Jing Hao, Jaemin Cho, Mohit Bansal, Jaehong Yoon
The paper presents a new approach to automatic audio description (AD) that treats the task as a constrained global optimization problem. It jointly decides what visual content is narratively important, when it can be spoken without overlapping dialogue, and how to phrase it within time limits. Using large language models for salience estimation and a mixed‑integer linear program for scheduling, the system outperforms prior methods on the REFRAMED benchmark, especially in temporal placement and narrative relevance.
By Igor Sterner, Mirella Lapata, Alex Lascarides, Frank Keller
arXiv:2608.21430v1 Announce Type: new
Abstract: Multimodal language models increasingly show promise for enabling the large-scale computational analysis of film, opening up new avenues for learning a...
By David Bamman, Kent K. Chang, Allison Cooper, Juishan Hsu, Reina Kushihashi, Madison Mar, Arnav Podichetty, Rachael Samberg, Ipek Nil Sancak, Yuhan Shao
arXiv:2608.28823v1 Announce Type: cross
Abstract: Artists coordinate human pose, illumination, and camera placement to convey narrative and emotion, but existing generative methods typically model th...
By Yunge Wen
The paper presents a method for generating 3D staging—human poses, lighting, and camera setup—directly from affective textual descriptions. It builds a dataset of 11,911 text–staging pairs derived from 2,328 figurative paintings, reconstructing SMPL bodies, estimating illumination, and recovering camera parameters. A flow‑matching transformer is trained to produce variable‑size scenes and multiple staging alternatives, achieving a 32.2% retrieval R@1 on held‑out prompts, outperforming a CLIP‑based baseline.
By Yunge Wen
arXiv:2608.30068v1 Announce Type: new
Abstract: Films communicate through deliberate creative choices, including lighting, color, composition, editing, dialogue, music, and sound. Humans naturally in...
By Kaishuu Shinozaki-Conefrey, Olivier Pascaud, Robin Courant, Xi Wang, Dimitris Samaras, Vicky Kalogeiton
SonicCaps is a large-scale audio captioning dataset featuring approximately 15 million captions paired with 700,000 audio clips, created using the Qwen3-Omni multimodal language model. The dataset emphasizes diversity by generating around 24 captions per clip through structured prompt engineering and few-shot generation, covering main descriptions, rephrased variants, and semantic tags. Human evaluations rate SonicCaps higher than existing datasets, and training CLAP models on it improves audio retrieval and zero-shot classification across public and commercial benchmarks.
By Zineb Lahrichi, Marc Ferras, Ga\"el Richard, Geoffroy Peeters
The paper introduces MObyGaze, a dataset of 20 films annotated by experts for multimodal objectification, covering 6072 segments across 43 hours of video. It defines objectification through a structured thesaurus of 5 sub‑constructs and 11 concepts spanning visual, speech, and audio modalities. The authors formulate learning tasks, explore label diversity strategies, and benchmark vision, text, and audio models to demonstrate the task’s feasibility.
By Julie Tores, Elisa Ancarani, Lucile Sassatelli, Hui-Yin Wu, Clement Bergman, Lea Andolfi, Victor Ecrement, Remy Sun, Frederic Precioso, Thierry Devars, Magali Guaresi, Virginie Julliard, Sarah Lecossais
arXiv:2608.29644v1 Announce Type: cross
Abstract: Attributing an artwork to an artist has traditionally relied on detailed visual observations and descriptions, known as stylistic analysis in art his...
By Marc S. Walton, Astrid Harth
Abstract4D is the largest dataset of abstract paintings, containing over 120,000 images with rich metadata and multi‑dimensional prompts that capture perceptual attributes such as form, color, texture, and composition. The dataset is annotated via a hybrid human–VLM pipeline to ensure quality and consistency. Using Abstract4D, the authors analyze the semantic structure of abstract art through large‑scale embedding visualization and establish benchmark tasks for classification, cross‑modal retrieval, and text‑to‑image generation to evaluate AI models’ perception and reproduction of abstract visual language.
By Haowei Zhang, Yuanpei Zhao, Ji-Zhe Zhou, Mao Li
arXiv:2605. 16972v2 Announce Type: replace-cross Abstract: Cultural heritage exhibitions often struggle to sustain attention and support reflective engagement.
By Jingjing Li, Zhi Liu, Xiyao Jin, Tatsuki Fushimi, Yoichi Ochiai
The paper introduces Redemption Score (RS), a multi‑modal evaluation framework for image captioning that combines three complementary signals: Mutual Information Divergence for global image‑text alignment, DINO‑based perceptual similarity of cycle‑generated images for visual grounding, and LLM text embeddings for contextual similarity to human references. RS fuses these signals to provide a more holistic assessment, achieving a Kendall‑τ of 58.42 on Flickr8k and outperforming most prior methods. The framework demonstrates consistent performance across Conceptual Captions and MS COCO, offering a robust evaluation that captures both visual accuracy and text quality.
By Ashim Dahal, Ankit Ghimire, Saydul Akbar Murad, Nick Rahimi