arXiv AI By Weijie Wu, Junbo Li, Lin Li, Jun Fang, Qingyang Hong

MMAC: A Massive Multi-dimensional Benchmark for Audio Captioning

Read the original on arXiv AI →

arXiv:2607. 27109v2 Announce Type: cross Abstract: With the development of audio large language models (AudioLLMs), audio captioning needs to move from brief descriptions toward open-ended and fine-grained free-form descriptions.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Computation and Language
Sep 25

An Evaluation Framework for Structured Audio Captions Validated by Controlled Perturbations

The paper introduces an evaluation framework for structured audio captions that separates acoustic and semantic aspects, such as timestamped sound event descriptions. It covers five axes—tag sets, descriptions, reasoning, numeric measurements, and spectral profiles—using large language model judges for semantics and deterministic metrics for temporal and acoustic features. Controlled perturbations validate that the metrics are robust to paraphrases but sensitive to real semantic and acoustic errors.

By Liang-Yuan Wu, Sripathi Sridhar, Mark Cartwright, Magdalena Fuentes
arXiv Computation and Language
Sep 3

SonicCaps: Large-Scale Diverse and Fine-Grained Captioning for Improved Audio-Retrieval

SonicCaps is a large-scale audio captioning dataset featuring approximately 15 million captions paired with 700,000 audio clips, created using the Qwen3-Omni multimodal language model. The dataset emphasizes diversity by generating around 24 captions per clip through structured prompt engineering and few-shot generation, covering main descriptions, rephrased variants, and semantic tags. Human evaluations rate SonicCaps higher than existing datasets, and training CLAP models on it improves audio retrieval and zero-shot classification across public and commercial benchmarks.

By Zineb Lahrichi, Marc Ferras, Ga\"el Richard, Geoffroy Peeters
arXiv Computer Vision
Sep 3

From Visual Cues to Spoken Narration: Rethinking Audio Description

The paper introduces Cue2Narrate, a two‑stage pipeline that jointly predicts what visual events to narrate and when to insert the narration in long, untrimmed movie clips. It uses a dual‑head audio‑visual localizer to identify visual cue and narration windows, followed by a LoRA‑adapted vision‑language model that generates concise audio descriptions, trained with a Description Ranking Loss. The authors also present the LongLSMDC benchmark, comprising up to 8‑minute clips, and show that Cue2Narrate outperforms video‑only and audio‑only baselines by 5–12 points in average mAP and improves AD generation over fine‑tuned base VLMs.

By Akshita Gupta, Aditya Arora, Federico Tombari, Marcus Rohrbach, Anna Rohrbach