Hugging Face Trending Papers

Drive-to-Music: Context-Aware Generative Audio for In-Vehicle Experiences

In-vehicle music can serve as an adaptive interface to enhance driver experience, attention, and well-being. We present Drive-to-Music, a context-aware system that generates music in real time from multimodal driving signals.

arXiv Computer Vision
Sep 3

Video Object Segmentation-Aware Audio Generation

The paper introduces a new task called video object segmentation‑aware audio generation, which conditions sound synthesis on object‑level segmentation maps. It presents SAGANet, a multimodal generative model that uses visual segmentation masks, video, and textual cues to produce controllable audio for musical instruments, offering fine‑grained, visually localized control. The authors also release the Segmented Music Solos dataset of instrument performance videos with segmentation information to support this task and demonstrate that SAGANet outperforms current state‑of‑the‑art methods in controllable, high‑fidelity Foley synthesis.

By Ilpo Viertola, Vladimir Iashin, Esa Rahtu
Hugging Face Trending Papers
Aug 6

Vorch-Omni: Multi-Task Orchestration of Sight and Sound

Recent advances in generative video modeling have enabled diverse generation, reference-based synthesis, extension, and editing, but existing approaches often rely on fragmented task-specific models. A general model must distinguish heterogeneous target, source, and reference signals to determine what to generate, preserve, or use as guidance, while reducing interference among tasks.

Hugging Face Trending Papers
Jun 24

From Sounds to Scenes: A Benchmark for Evaluating Context-Aware Auditory Scene Understanding in Large Audio Language Models

Recent Large Audio Language Models (LALMs) have achieved remarkable progress in audio perceptual tasks across individual acoustic layers, including speech, sound, and music. However, existing benchmarks predominantly evaluate these layers in isolation, overlooking the complex contextual relationships that arise when multiple acoustic sources co-occur in real-world auditory scenes.

Hugging Face Trending Papers
Jul 15

Bring Music The Horizon: Music-Driven 360$^\circ$ Video Generation

Music visualization offers a powerful way to enhance listeners' understanding and experience of music by translating auditory signals into visual forms. However, most existing approaches either rely heavily on lyrics or generate flat, non-immersive videos similar to conventional music videos, which limits their ability to convey the emotional dynamics of music and provide an immersive listening experience.