arXiv AI

CANVAS: Captioning Art with Narrative Visual-Audio AI Systems

arXiv:2606. 09846v1 Announce Type: cross Abstract: Visual art remains largely inaccessible to blind and low-vision (BLV) audiences due to brief or absent alt-text, which rarely conveys the sensory, spatial, or emotional qualities of an artwork.

arXiv AI
Jun 30

MuseBench: Benchmarking Intent-Level Audiovisual Arts Understanding in MLLMs

arXiv:2606. 30026v1 Announce Type: cross Abstract: Audiovisual arts encompass diverse creative disciplines, including cinema, visual arts, stage performance, and game design, where artistic meaning arises from deliberate combinations of visual, auditory, and narrative elements (e.

By Yuxuan Fan, Gyusik Seo, Jing Hao, Jaemin Cho, Mohit Bansal, Jaehong Yoon
arXiv Computation and Language
Sep 25

What, When, and How: Audio Description as Constrained Global Optimization

The paper presents a new approach to automatic audio description (AD) that treats the task as a constrained global optimization problem. It jointly decides what visual content is narratively important, when it can be spoken without overlapping dialogue, and how to phrase it within time limits. Using large language models for salience estimation and a mixed‑integer linear program for scheduling, the system outperforms prior methods on the REFRAMED benchmark, especially in temporal placement and narrative relevance.

By Igor Sterner, Mirella Lapata, Alex Lascarides, Frank Keller
arXiv AI
Sep 16

Text-Driven Artistic Staging: 3D Posing, Lighting, and Camera References from Paintings

The paper presents a method for generating 3D staging—human poses, lighting, and camera setup—directly from affective textual descriptions. It builds a dataset of 11,911 text–staging pairs derived from 2,328 figurative paintings, reconstructing SMPL bodies, estimating illumination, and recovering camera parameters. A flow‑matching transformer is trained to produce variable‑size scenes and multiple staging alternatives, achieving a 32.2% retrieval R@1 on held‑out prompts, outperforming a CLIP‑based baseline.

By Yunge Wen
arXiv Computation and Language
Sep 3

SonicCaps: Large-Scale Diverse and Fine-Grained Captioning for Improved Audio-Retrieval

SonicCaps is a large-scale audio captioning dataset featuring approximately 15 million captions paired with 700,000 audio clips, created using the Qwen3-Omni multimodal language model. The dataset emphasizes diversity by generating around 24 captions per clip through structured prompt engineering and few-shot generation, covering main descriptions, rephrased variants, and semantic tags. Human evaluations rate SonicCaps higher than existing datasets, and training CLAP models on it improves audio retrieval and zero-shot classification across public and commercial benchmarks.

By Zineb Lahrichi, Marc Ferras, Ga\"el Richard, Geoffroy Peeters
arXiv Computer Vision
Aug 27

MObyGaze: a film dataset of multimodal objectification densely annotated by experts

The paper introduces MObyGaze, a dataset of 20 films annotated by experts for multimodal objectification, covering 6072 segments across 43 hours of video. It defines objectification through a structured thesaurus of 5 sub‑constructs and 11 concepts spanning visual, speech, and audio modalities. The authors formulate learning tasks, explore label diversity strategies, and benchmark vision, text, and audio models to demonstrate the task’s feasibility.

By Julie Tores, Elisa Ancarani, Lucile Sassatelli, Hui-Yin Wu, Clement Bergman, Lea Andolfi, Victor Ecrement, Remy Sun, Frederic Precioso, Thierry Devars, Magali Guaresi, Virginie Julliard, Sarah Lecossais
arXiv Computer Vision
Aug 31

Abstract4D: A Large-Scale Dataset and Framework for Understanding the Visual Language of Abstract Art

Abstract4D is the largest dataset of abstract paintings, containing over 120,000 images with rich metadata and multi‑dimensional prompts that capture perceptual attributes such as form, color, texture, and composition. The dataset is annotated via a hybrid human–VLM pipeline to ensure quality and consistency. Using Abstract4D, the authors analyze the semantic structure of abstract art through large‑scale embedding visualization and establish benchmark tasks for classification, cross‑modal retrieval, and text‑to‑image generation to evaluate AI models’ perception and reproduction of abstract visual language.

By Haowei Zhang, Yuanpei Zhao, Ji-Zhe Zhou, Mao Li
arXiv Computation and Language
Sep 18

Redemption Score: A Multi-Modal Evaluation Framework for Image Captioning via Distributional, Perceptual, and Linguistic Signal Triangulation

The paper introduces Redemption Score (RS), a multi‑modal evaluation framework for image captioning that combines three complementary signals: Mutual Information Divergence for global image‑text alignment, DINO‑based perceptual similarity of cycle‑generated images for visual grounding, and LLM text embeddings for contextual similarity to human references. RS fuses these signals to provide a more holistic assessment, achieving a Kendall‑τ of 58.42 on Flickr8k and outperforming most prior methods. The framework demonstrates consistent performance across Conceptual Captions and MS COCO, offering a robust evaluation that captures both visual accuracy and text quality.

By Ashim Dahal, Ankit Ghimire, Saydul Akbar Murad, Nick Rahimi