The paper investigates whether joint language‑audio embedding models encode human perceptual timbre semantics. It evaluates several state‑of‑the‑art models, finding that LAION‑CLAP aligns best with human‑perceived timbre across instrumental sounds and descriptor‑conditioned audio effects, yet the overall alignment remains limited. The study also notes that reverb‑induced timbre semantics are more consistently captured than equalization‑induced ones.
By Qixin Deng, Bryan Pardo, Thrasyvoulos N Pappas
arXiv:2606. 06615v1 Announce Type: cross Abstract: Retrieving music using natural language descriptions has improved with contrastive audio-text models such as CLAP, but current systems remain limited to coarse semantic queries.
By Nishit Anand, Ashish Seth, Sreyan Ghosh, Dinesh Manocha, Ramani Duraiswami
TTM-Bench is a framework designed to benchmark text-to-music systems by establishing a common protocol for reproducible evaluation. It measures performance along two axes: musical-content alignment—assessed through semantic, genre, and musical-descriptor agreement scores against a shared musical specification—and computational efficiency, which includes generation latency, real-time factor, resource usage for local models, and cost for hosted services. A preliminary case study using TTM-Bench shows that higher alignment does not necessarily mean lower computational demands, underscoring the need for distinct, interpretable metrics.
By Giorgia Adorni, Michela Papandrea, Battista Rimoldi, Tiziano Leidi
The paper introduces an evaluation framework for structured audio captions that separates acoustic and semantic aspects, such as timestamped sound event descriptions. It covers five axes—tag sets, descriptions, reasoning, numeric measurements, and spectral profiles—using large language model judges for semantics and deterministic metrics for temporal and acoustic features. Controlled perturbations validate that the metrics are robust to paraphrases but sensitive to real semantic and acoustic errors.
By Liang-Yuan Wu, Sripathi Sridhar, Mark Cartwright, Magdalena Fuentes
SonicCaps is a large-scale audio captioning dataset featuring approximately 15 million captions paired with 700,000 audio clips, created using the Qwen3-Omni multimodal language model. The dataset emphasizes diversity by generating around 24 captions per clip through structured prompt engineering and few-shot generation, covering main descriptions, rephrased variants, and semantic tags. Human evaluations rate SonicCaps higher than existing datasets, and training CLAP models on it improves audio retrieval and zero-shot classification across public and commercial benchmarks.
By Zineb Lahrichi, Marc Ferras, Ga\"el Richard, Geoffroy Peeters
MUUNRiver-Bench is a diagnostic benchmark for music retrieval that uses natural‑language instructions to define relevance for reference‑audio queries. It contains 3,440 tracks across 13 genres and 116 sub‑genres and covers seven tasks such as similar‑music, style‑preserving lyric‑rewriting, cover, and segment retrieval. Experiments with six models in eight configurations show that acoustic encoders favor local identity while text‑aligned encoders favor semantic relations, and that instruction‑aware and audio‑text fusion systems do not consistently outperform their backbones.
By Zhancheng Guo, Congren Dai, Shangda Wu, Jianhuai Hu, Danni Zhao, Xiaobing Li, Maosong Sun