arXiv:2508.03448v4 Announce Type: replace-cross
Abstract: Music recordings often suffer from audio quality issues such as excessive reverberation, distortion, clipping, tonal imbalances, and a narrow...
By Jan Melechovsky, Ambuj Mehrish, Abhinaba Roy, Dorien Herremans
Spot, Separate, and Enhance (SSE) is a multimodal, user‑guided generative model for audio remixing and enhancement. It rebalances audio, removes unwanted sources, and reduces reverberation using video and textual guidance. The authors introduce the DegradedMix dataset and adopt generative evaluation metrics, showing SSE outperforms existing baselines in controllability and remixing quality.
By Ilpo Viertola, Giulio Cengarle, Gouthaman KV, Daniel Arteaga, Lie Lu
arXiv:2608. 04142v1 Announce Type: cross Abstract: Existing reference-free methods for evaluating music perceptual quality alleviate the need for paired noisy-clean data, but they still rely on a background set, which is used to compute aggregated statistics of clean audio samples.
By Alon Ziv, Harel Pogoda, Yossi Adi
Generalizable Audio-to-Score (A2S) transcription is fundamentally constrained by the severe scarcity of high-quality, real-world paired data. Relying solely on existing human-annotated datasets often...
arXiv:2606. 13626v2 Announce Type: replace-cross Abstract: We study generative modeling of Bach-style symbolic piano music using a shared MIDI corpus and three model families: autoregressive LSTMs with attention, latent-variable models including recurrent VAEs and vector-quantized VAEs, and generative adversarial networks.
By Dezhi Yu, Kyuil Lee, Yongkang Huang
arXiv:2506. 14293v4 Announce Type: replace-cross Abstract: We present Sleeping-DISCO 9M, a large-scale pre-training dataset for music and song.
By Tawsif Ahmed, Andrej Radonjic, Gollam Rabby