Vocal Music under Phoneme-Conditional Analysis
arXiv:2608.30823v1 Announce Type: cross Abstract: The vocal music of each language carries a distinctive sonic identity, even without instrumental accompaniment. We ask whether these differences are...
arXiv:2608.30823v1 Announce Type: cross Abstract: The vocal music of each language carries a distinctive sonic identity, even without instrumental accompaniment. We ask whether these differences are...
arXiv:2607. 26698v1 Announce Type: cross Abstract: Cover song generation (CSG) should preserve the melodic and linguistic content of a reference song while recreating the remaining musical components.
arXiv:2606. 26451v1 Announce Type: cross Abstract: Automatic singing quality assessment (SQA) requires evaluating lyrical correctness and musical fidelity while handling expressive variations.
arXiv:2607. 05902v1 Announce Type: cross Abstract: Chamber music, as a highly precise multi-part interactive system, contains a logic of "role assignment and dynamic interaction" that provides an extremely valuable blueprint for exploring human-computer collaborative composition paradigms.
arXiv:2606. 05852v1 Announce Type: cross Abstract: Text-to-speech (TTS) and singing voice synthesis (SVS) both aim to generate human vocal audio from symbolic inputs, but they impose different requirements on the generation process.
The paper evaluates whether music‑text models truly capture fine‑grained musical meaning by introducing attribute‑swap perturbations that exchange properties such as timbre or order between instruments in a caption. Four contrastive models and one large audio‑language model were tested to see if they would score higher on the original caption than on the perturbed one. The results show that none of the contrastive models reliably distinguish the captions, and the audio‑language model’s advantage stems mainly from language priors, indicating that CLAP scores behave like a bag‑of‑words and fail to reflect attribute bindings.
The paper presents the first controlled study of Whisper adaptation for Greek Automatic Lyric Transcription (ALT), addressing challenges such as melodic variability, rhythmic irregularity, and accompaniment interference. It explores model scaling, multitask training with transcribe-translate ratios, and a two‑stage speech‑to‑singing adaptation, while curating a segment‑level aligned singing dataset from the Greek Audio Dataset. Results show that larger models consistently improve performance, multitask learning benefits smaller models, and the two‑stage adaptation achieves a 27.2% WER, establishing the first Greek ALT benchmark.
arXiv:2603. 28378v2 Announce Type: replace-cross Abstract: We present the first systematic Membership Inference Attack (MIA) evaluation of LALMs.
The paper investigates whether audio language models encode phonetic features similarly when processing spoken versus written input. By comparing mean representations of minimal phoneme pairs across six models, seven features, and 15 languages, the study finds that only voicing in two Qwen2.5-Omni models shows a significant shared direction, and that the model family—not size—determines feature representation. The analysis uses cosine similarity against a random-pair reference to assess alignment across modalities.
Neural TTS systems can sound natural across languages, but naturalness does not guarantee the preservation of sound contrasts that distinguish words from their grammatical forms. Standard metrics like MOS do not test for this.
Project Qualia investigates whether experiential similarity between songs can be extracted from listening behavior. Using 1.29 billion scrobbles from 9,396 users, the authors trained a Word2Vec model (Song2Vec) on session data, then applied an artist‑residual procedure to isolate artist‑independent signals. The residual embeddings still contained strong cross‑artist similarity, forming coherent genre and era clusters, demonstrating that experiential structure exists beyond artist identity.
arXiv:2506. 14293v4 Announce Type: replace-cross Abstract: We present Sleeping-DISCO 9M, a large-scale pre-training dataset for music and song.