arXiv Machine Learning

Listening Like a Judge: A Music-Aware Framework for Automatic Singing Performance Evaluation

arXiv:2606. 26451v1 Announce Type: cross Abstract: Automatic singing quality assessment (SQA) requires evaluating lyrical correctness and musical fidelity while handling expressive variations.

arXiv AI
Sep 15

DuoTok: Source-Aware Dual-Track Music Tokenization for Vocal-Accompaniment Generation

DuoTok is a source‑aware dual‑track music tokenizer designed for vocal‑accompaniment generation. It first learns a semantic audio representation via self‑supervised pretraining, then refines source‑aware structure with feature‑replacement noise and multi‑task supervision (spectral reconstruction, source separation regularization, and an ASR head for lyric alignment). The encoder is frozen and hard‑routed codebooks for vocals and accompaniment are learned, while a diffusion decoder restores fine acoustic detail from the discrete tokens, achieving a favorable predictability‑fidelity trade‑off at ultra‑low bitrate across public benchmarks.

By Rui Lin, Zhiyue Wu, Jiahe Lei, Kangdi Wang, Weixiong Chen, Junyu Dai, Tao Jiang
arXiv AI
Sep 16

MUUNRiver-Bench: Diagnosing Relation-Dependent Music Retrieval with Multimodal Instructions

MUUNRiver-Bench is a diagnostic benchmark for music retrieval that uses natural‑language instructions to define relevance for reference‑audio queries. It contains 3,440 tracks across 13 genres and 116 sub‑genres and covers seven tasks such as similar‑music, style‑preserving lyric‑rewriting, cover, and segment retrieval. Experiments with six models in eight configurations show that acoustic encoders favor local identity while text‑aligned encoders favor semantic relations, and that instruction‑aware and audio‑text fusion systems do not consistently outperform their backbones.

By Zhancheng Guo, Congren Dai, Shangda Wu, Jianhuai Hu, Danni Zhao, Xiaobing Li, Maosong Sun
Hugging Face Trending Papers
Jul 22

Pushing the Frontier of Full-Song Generation: Hierarchical Autoregressive Planning Meets Flow-Matching Rendering

In this report, we present a unified song generation framework capable of producing high-quality full-length music from lyrics, text descriptions, and musical attributes. The proposed framework supports three tasks: Lyrics-to-Song Generation, which generates complete songs from text descriptions, lyrics, and musical attributes; Instrumental Music Generation, which creates music without vocals; and Cover Song Generation, which reinterprets existing songs with different styles while preserving their melodic content.

arXiv Computation and Language
Sep 11

Automatic Lyric Transcription for Greek Songs: Scaling and Task Composition Effects in Whisper Adaptation

The paper presents the first controlled study of Whisper adaptation for Greek Automatic Lyric Transcription (ALT), addressing challenges such as melodic variability, rhythmic irregularity, and accompaniment interference. It explores model scaling, multitask training with transcribe-translate ratios, and a two‑stage speech‑to‑singing adaptation, while curating a segment‑level aligned singing dataset from the Greek Audio Dataset. Results show that larger models consistently improve performance, multitask learning benefits smaller models, and the two‑stage adaptation achieves a 27.2% WER, establishing the first Greek ALT benchmark.

By Maria Frangiadaki, Dimitrios Damianos, Kosmas Kritsis, Vassilis Katsouros
arXiv AI
Jul 23

Pushing the Frontier of Full-Song Generation: Hierarchical Autoregressive Planning Meets Flow-Matching Rendering

arXiv:2607. 20253v1 Announce Type: cross Abstract: In this report, we present a unified song generation framework capable of producing high-quality full-length music from lyrics, text descriptions, and musical attributes.

By Junyu Dai, Xinyue Fan, Weiqin Li, Xiangang Li, Yunjia Li, Bin Ma, Yukun Ma, Chongjia Ni, Yufei Shi, Haoxu Wang, Menglin Wu, Jianwei Yu, Huaicheng Zhang, Han Zhao, Shengkui Zhao, Haina Zhu