Hugging Face Trending Papers

Spot, Separate, and Enhance: Fully Generative Approach for Audio Mixing

Spot, Separate, and Enhance (SSE) is the first multimodal, user‑guided generative model for audio remixing and enhancement. It rebalances audio, removes unwanted sources, and reduces reverberation in video content, guided by both video and textual descriptions. The authors introduce the DegradedMix dataset, built on MuddyMix, and use generative‑model evaluation metrics to demonstrate SSE’s superior controllability and remixing quality compared to existing baselines.

arXiv AI
3d ago

Spot, Separate, and Enhance: Fully Generative Approach for Audio Mixing

Spot, Separate, and Enhance (SSE) is a multimodal, user‑guided generative model for audio remixing and enhancement. It rebalances audio, removes unwanted sources, and reduces reverberation using video and textual guidance. The authors introduce the DegradedMix dataset and adopt generative evaluation metrics, showing SSE outperforms existing baselines in controllability and remixing quality.

By Ilpo Viertola, Giulio Cengarle, Gouthaman KV, Daniel Arteaga, Lie Lu
arXiv Machine Learning
Jul 28

Music-Source-Separation-Training (MSST): A Unified Framework for Training and Evaluating Music Demixing Models

arXiv:2607. 23395v1 Announce Type: cross Abstract: Music Source Separation (MSS), the task of recovering individual sound components (stems) from a polyphonic mixture, is central to applications ranging from karaoke and remixing to audio restoration and content production.

By Roman Solovyev, Ilya Kiselev, Alexander Stempkovskiy, Tatiana Gabruseva
arXiv Computer Vision
Sep 3

Video Object Segmentation-Aware Audio Generation

The paper introduces a new task called video object segmentation‑aware audio generation, which conditions sound synthesis on object‑level segmentation maps. It presents SAGANet, a multimodal generative model that uses visual segmentation masks, video, and textual cues to produce controllable audio for musical instruments, offering fine‑grained, visually localized control. The authors also release the Segmented Music Solos dataset of instrument performance videos with segmentation information to support this task and demonstrate that SAGANet outperforms current state‑of‑the‑art methods in controllable, high‑fidelity Foley synthesis.

By Ilpo Viertola, Vladimir Iashin, Esa Rahtu
arXiv AI
Jul 23

Pushing the Frontier of Full-Song Generation: Hierarchical Autoregressive Planning Meets Flow-Matching Rendering

arXiv:2607. 20253v1 Announce Type: cross Abstract: In this report, we present a unified song generation framework capable of producing high-quality full-length music from lyrics, text descriptions, and musical attributes.

By Junyu Dai, Xinyue Fan, Weiqin Li, Xiangang Li, Yunjia Li, Bin Ma, Yukun Ma, Chongjia Ni, Yufei Shi, Haoxu Wang, Menglin Wu, Jianwei Yu, Huaicheng Zhang, Han Zhao, Shengkui Zhao, Haina Zhu