arXiv AI By Chun-Yi Kuan, Hung-yi Lee

The Sound of Absence: Audio-Language Embedding Models Struggle with Negation

Read the original on arXiv AI →

arXiv:2607. 12290v1 Announce Type: cross Abstract: Audio-language embedding models such as CLAP are widely evaluated on matching present sound events, but rarely on negation.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Aug 26

Do Joint Language-Audio Embeddings Encode Perceptual Timbre Semantics?

The paper investigates whether joint language‑audio embedding models encode human perceptual timbre semantics. It evaluates several state‑of‑the‑art models, finding that LAION‑CLAP aligns best with human‑perceived timbre across instrumental sounds and descriptor‑conditioned audio effects, yet the overall alignment remains limited. The study also notes that reverb‑induced timbre semantics are more consistently captured than equalization‑induced ones.

By Qixin Deng, Bryan Pardo, Thrasyvoulos N Pappas
arXiv Computation and Language
Sep 3

SonicCaps: Large-Scale Diverse and Fine-Grained Captioning for Improved Audio-Retrieval

SonicCaps is a large-scale audio captioning dataset featuring approximately 15 million captions paired with 700,000 audio clips, created using the Qwen3-Omni multimodal language model. The dataset emphasizes diversity by generating around 24 captions per clip through structured prompt engineering and few-shot generation, covering main descriptions, rephrased variants, and semantic tags. Human evaluations rate SonicCaps higher than existing datasets, and training CLAP models on it improves audio retrieval and zero-shot classification across public and commercial benchmarks.

By Zineb Lahrichi, Marc Ferras, Ga\"el Richard, Geoffroy Peeters
arXiv Computation and Language
Sep 25

An Evaluation Framework for Structured Audio Captions Validated by Controlled Perturbations

The paper introduces an evaluation framework for structured audio captions that separates acoustic and semantic aspects, such as timestamped sound event descriptions. It covers five axes—tag sets, descriptions, reasoning, numeric measurements, and spectral profiles—using large language model judges for semantics and deterministic metrics for temporal and acoustic features. Controlled perturbations validate that the metrics are robust to paraphrases but sensitive to real semantic and acoustic errors.

By Liang-Yuan Wu, Sripathi Sridhar, Mark Cartwright, Magdalena Fuentes
arXiv AI
Aug 5

From Generator to Embedder: Harnessing Innate Abilities of Multimodal LLMs via Building Zero-Shot Discriminative Embedding Model

arXiv:2508. 00955v3 Announce Type: replace-cross Abstract: Adapting generative Multimodal Large Language Models (MLLMs) into universal embedding models typically demands resource-intensive contrastive pre-training, while traditional hard negative mining methods suffer from severe false negative contamination.

By Yeong-Joon Ju, Seong-Whan Lee