Spot, Separate, and Enhance (SSE) is the first multimodal, user‑guided generative model for audio remixing and enhancement. It rebalances audio, removes unwanted sources, and reduces reverberation in video content, guided by both video and textual descriptions. The authors introduce the DegradedMix dataset, built on MuddyMix, and use generative‑model evaluation metrics to demonstrate SSE’s superior controllability and remixing quality compared to existing baselines.
arXiv:2508.03448v4 Announce Type: replace-cross
Abstract: Music recordings often suffer from audio quality issues such as excessive reverberation, distortion, clipping, tonal imbalances, and a narrow...
By Jan Melechovsky, Ambuj Mehrish, Abhinaba Roy, Dorien Herremans
arXiv:2602. 03762v4 Announce Type: replace-cross Abstract: Visually-guided acoustic highlighting seeks to rebalance audio in alignment with the accompanying video, creating a coherent audio-visual experience.
By Hugo Malard, Gael Le Lan, Daniel Wong, David Lou Alon, Yi-Chiao Wu, Sanjeel Parekh
arXiv:2608. 04142v1 Announce Type: cross Abstract: Existing reference-free methods for evaluating music perceptual quality alleviate the need for paired noisy-clean data, but they still rely on a background set, which is used to compute aggregated statistics of clean audio samples.
By Alon Ziv, Harel Pogoda, Yossi Adi
arXiv:2603. 09234v2 Announce Type: cross Abstract: Achieving high perceptual quality without hallucination remains a challenge in generative speech enhancement (SE).
By Xiaobin Rong, Jun Gao, Zheng Wang, Mansur Yesilbursa, Kamil Wojcicki, Jing Lu
arXiv:2506. 20995v4 Announce Type: replace-cross Abstract: We propose a step-by-step video-to-audio (V2A) generation method that provides finer control over the generation process and more realistic audio synthesis.
By Akio Hayakawa, Masato Ishii, Takashi Shibuya, Yuki Mitsufuji