arXiv AI

AVSD-Scenes: A Dataset for Audio-Visual Description of Urban Scenes

AVSD-Scenes is a new dataset of 12,291 audio‑visual scene descriptions for urban environments, built from the TAU Urban Audio‑Visual Scenes dataset. The descriptions are generated by first creating modality‑specific text with Qwen2‑Audio‑7B and Qwen2.5‑VL‑7B, then merging them with large language models (Qwen3‑14B, Mistral‑Small‑3.2‑24B‑Instruct‑2506, Gemma‑3‑27B‑it) to produce multimodal narratives that combine auditory and visual cues. Benchmarks show that these multimodal descriptions improve semantic alignment, cross‑modal retrieval, and scene classification accuracy (up to 95.4%) while remaining discriminative even without explicit scene labels.

arXiv Computation and Language
Sep 3

SonicCaps: Large-Scale Diverse and Fine-Grained Captioning for Improved Audio-Retrieval

SonicCaps is a large-scale audio captioning dataset featuring approximately 15 million captions paired with 700,000 audio clips, created using the Qwen3-Omni multimodal language model. The dataset emphasizes diversity by generating around 24 captions per clip through structured prompt engineering and few-shot generation, covering main descriptions, rephrased variants, and semantic tags. Human evaluations rate SonicCaps higher than existing datasets, and training CLAP models on it improves audio retrieval and zero-shot classification across public and commercial benchmarks.

By Zineb Lahrichi, Marc Ferras, Ga\"el Richard, Geoffroy Peeters
arXiv Machine Learning
Jul 14

Empowering Long-form Omni-modal Understanding with Robust Audio Perception

arXiv:2607. 10299v1 Announce Type: new Abstract: Recent advances in large-scale multimodal models have drivenremarkable progress in vision-language tasks; however, comprehensiveomni-modal understanding remains under-explored, largely due to thescarcity of datasets with rich, explicitly aligned auditory cues.

By Kaiying Yan, Luoyi Sun, Xiao Zhou, Weidi Xie
arXiv Computer Vision
Sep 3

From Visual Cues to Spoken Narration: Rethinking Audio Description

The paper introduces Cue2Narrate, a two‑stage pipeline that jointly predicts what visual events to narrate and when to insert the narration in long, untrimmed movie clips. It uses a dual‑head audio‑visual localizer to identify visual cue and narration windows, followed by a LoRA‑adapted vision‑language model that generates concise audio descriptions, trained with a Description Ranking Loss. The authors also present the LongLSMDC benchmark, comprising up to 8‑minute clips, and show that Cue2Narrate outperforms video‑only and audio‑only baselines by 5–12 points in average mAP and improves AD generation over fine‑tuned base VLMs.

By Akshita Gupta, Aditya Arora, Federico Tombari, Marcus Rohrbach, Anna Rohrbach
arXiv Computer Vision
Aug 27

Asymmetric Cross-Modal Fine-Grained Visual Categorization: ACF-Net and the BirdPro Benchmark

The paper introduces ACF-Net, an optical flow‑guided framework for asymmetric audio‑visual fine‑grained visual categorization (FGVC), addressing challenges where video and audio are not strictly synchronized or matched. ACF-Net comprises Optical Flow‑Guided Motion (OFGM) to capture motion‑sensitive visual cues and suppress background noise, and Asymmetric Cross‑Modal Adaptive Fusion (ACAF) to estimate modality reliability and perform uncertainty‑aware fusion. The authors also present BirdPro, a new bird‑oriented audio‑visual benchmark with 1,919 audio recordings and 11,965 videos across 194 species, and report that ACF‑Net outperforms baselines by 2.97% in fused and 1.92% in mismatched settings.

By Bohan Deng, Shuo Ye, Zitong Yu
arXiv AI
2d ago

SONIC-O1: A Real-World Benchmark for Evaluating Multimodal Large Language Models on Audio-Video Understanding

SONIC‑O1 is a new benchmark designed to evaluate multimodal large language models on audio‑video understanding. It contains 60 hours of 231 clips across 13 real‑world conversational domains, with 4,958 human‑verified annotations and demographic metadata. The benchmark tests open‑ended summarization, multiple‑choice question answering, and temporally grounded reasoning, revealing performance gaps between model families and across demographic groups.

By Ahmed Y. Radwan, Christos Emmanouilidis, Hina Tabassum, Deval Pandya, Shaina Raza
arXiv Computer Vision
Aug 27

MObyGaze: a film dataset of multimodal objectification densely annotated by experts

The paper introduces MObyGaze, a dataset of 20 films annotated by experts for multimodal objectification, covering 6072 segments across 43 hours of video. It defines objectification through a structured thesaurus of 5 sub‑constructs and 11 concepts spanning visual, speech, and audio modalities. The authors formulate learning tasks, explore label diversity strategies, and benchmark vision, text, and audio models to demonstrate the task’s feasibility.

By Julie Tores, Elisa Ancarani, Lucile Sassatelli, Hui-Yin Wu, Clement Bergman, Lea Andolfi, Victor Ecrement, Remy Sun, Frederic Precioso, Thierry Devars, Magali Guaresi, Virginie Julliard, Sarah Lecossais