arXiv AI

MUUNRiver-Bench: Diagnosing Relation-Dependent Music Retrieval with Multimodal Instructions

MUUNRiver-Bench is a diagnostic benchmark for music retrieval that uses natural‑language instructions to define relevance for reference‑audio queries. It contains 3,440 tracks across 13 genres and 116 sub‑genres and covers seven tasks such as similar‑music, style‑preserving lyric‑rewriting, cover, and segment retrieval. Experiments with six models in eight configurations show that acoustic encoders favor local identity while text‑aligned encoders favor semantic relations, and that instruction‑aware and audio‑text fusion systems do not consistently outperform their backbones.

arXiv AI
Jun 8

FIGMA: Towards FIne-Grained Music retrievAl

arXiv:2606. 06615v1 Announce Type: cross Abstract: Retrieving music using natural language descriptions has improved with contrastive audio-text models such as CLAP, but current systems remain limited to coarse semantic queries.

By Nishit Anand, Ashish Seth, Sreyan Ghosh, Dinesh Manocha, Ramani Duraiswami
arXiv AI
Jun 12

CMI-RewardBench: Evaluating Music Reward Models with Compositional Multimodal Instruction

arXiv:2603. 00610v3 Announce Type: replace-cross Abstract: While music generation models have evolved to handle complex multimodal inputs mixing text, lyrics, and reference audio, evaluation mechanisms have lagged behind.

By Yinghao Ma, Haiwen Xia, Hewei Gao, Weixiong Chen, Yuxin Ye, Yuchen Yang, Sungkyun Chang, Mingshuo Ding, Yizhi Li, Ruibin Yuan, Simon Dixon, Emmanouil Benetos
arXiv Machine Learning
Sep 11

Project Qualia: Recovering Experiential Music Structure from Session Co-occurrence Data

Project Qualia investigates whether experiential similarity between songs can be extracted from listening behavior. Using 1.29 billion scrobbles from 9,396 users, the authors trained a Word2Vec model (Song2Vec) on session data, then applied an artist‑residual procedure to isolate artist‑independent signals. The residual embeddings still contained strong cross‑artist similarity, forming coherent genre and era clusters, demonstrating that experiential structure exists beyond artist identity.

By Nizam Mohammed, Abu B. S. Rahman, Dimuthu D. K. Arachchige
arXiv AI
Aug 26

Do Joint Language-Audio Embeddings Encode Perceptual Timbre Semantics?

The paper investigates whether joint language‑audio embedding models encode human perceptual timbre semantics. It evaluates several state‑of‑the‑art models, finding that LAION‑CLAP aligns best with human‑perceived timbre across instrumental sounds and descriptor‑conditioned audio effects, yet the overall alignment remains limited. The study also notes that reverb‑induced timbre semantics are more consistently captured than equalization‑induced ones.

By Qixin Deng, Bryan Pardo, Thrasyvoulos N Pappas