Hugging Face Trending Papers

Vocal Music under Phoneme-Conditional Analysis

arXiv Computation and Language
Sep 1

Vocal Music under Phoneme-Conditional Analysis

arXiv:2608.30823v1 Announce Type: cross Abstract: The vocal music of each language carries a distinctive sonic identity, even without instrumental accompaniment. We ask whether these differences are...

By Hayoon Kim, Kyogu Lee
arXiv AI
Jul 8

From Textural Counterpoint to Feature Encoding: A Multi-Dimensional Machine Representation Study of Haydn's "The Lark" Integrating Electroacoustic Analysis

arXiv:2607. 05902v1 Announce Type: cross Abstract: Chamber music, as a highly precise multi-part interactive system, contains a logic of "role assignment and dynamic interaction" that provides an extremely valuable blueprint for exploring human-computer collaborative composition paradigms.

By Yakun Liu, Zhiyu Jin, Hai Luan, Dong Liu, Xiaonan Li
arXiv Computation and Language
3d ago

Don't CLAP: Are Music-Text Models Bag-of-Words?

The paper evaluates whether music‑text models truly capture fine‑grained musical meaning by introducing attribute‑swap perturbations that exchange properties such as timbre or order between instruments in a caption. Four contrastive models and one large audio‑language model were tested to see if they would score higher on the original caption than on the perturbed one. The results show that none of the contrastive models reliably distinguish the captions, and the audio‑language model’s advantage stems mainly from language priors, indicating that CLAP scores behave like a bag‑of‑words and fail to reflect attribute bindings.

By Yuan-Chiao Cheng, Alexander Lerch
arXiv Computation and Language
Sep 11

Automatic Lyric Transcription for Greek Songs: Scaling and Task Composition Effects in Whisper Adaptation

The paper presents the first controlled study of Whisper adaptation for Greek Automatic Lyric Transcription (ALT), addressing challenges such as melodic variability, rhythmic irregularity, and accompaniment interference. It explores model scaling, multitask training with transcribe-translate ratios, and a two‑stage speech‑to‑singing adaptation, while curating a segment‑level aligned singing dataset from the Greek Audio Dataset. Results show that larger models consistently improve performance, multitask learning benefits smaller models, and the two‑stage adaptation achieves a 27.2% WER, establishing the first Greek ALT benchmark.

By Maria Frangiadaki, Dimitrios Damianos, Kosmas Kritsis, Vassilis Katsouros
Hugging Face Trending Papers
Sep 24

Do Audio Language Models Hear and Read Distinctive Features Alike?

The paper investigates whether audio language models encode phonetic features similarly when processing spoken versus written input. By comparing mean representations of minimal phoneme pairs across six models, seven features, and 15 languages, the study finds that only voicing in two Qwen2.5-Omni models shows a significant shared direction, and that the model family—not size—determines feature representation. The analysis uses cosine similarity against a random-pair reference to assess alignment across modalities.

arXiv Machine Learning
Sep 11

Project Qualia: Recovering Experiential Music Structure from Session Co-occurrence Data

Project Qualia investigates whether experiential similarity between songs can be extracted from listening behavior. Using 1.29 billion scrobbles from 9,396 users, the authors trained a Word2Vec model (Song2Vec) on session data, then applied an artist‑residual procedure to isolate artist‑independent signals. The residual embeddings still contained strong cross‑artist similarity, forming coherent genre and era clusters, demonstrating that experiential structure exists beyond artist identity.

By Nizam Mohammed, Abu B. S. Rahman, Dimuthu D. K. Arachchige