Drive-to-Music: Context-Aware Generative Audio for In-Vehicle Experiences
arXiv:2608. 12615v1 Announce Type: cross Abstract: In-vehicle music can serve as an adaptive interface to enhance driver experience, attention, and well-being.
In-vehicle music can serve as an adaptive interface to enhance driver experience, attention, and well-being. We present Drive-to-Music, a context-aware system that generates music in real time from multimodal driving signals.
arXiv:2608. 12615v1 Announce Type: cross Abstract: In-vehicle music can serve as an adaptive interface to enhance driver experience, attention, and well-being.
arXiv:2608. 03742v1 Announce Type: cross Abstract: Sound effects play a crucial role in conveying actions, events, and environmental cues across digital applications, often requiring a high degree of variation and contextual adaptability.
arXiv:2606. 24307v1 Announce Type: cross Abstract: Interactive music and live performance relies on real-time human expression, but modern generative music AI remains largely absent from this domain due to its prohibitive inference latency and offline rendering paradigm.
The paper introduces a new task called video object segmentation‑aware audio generation, which conditions sound synthesis on object‑level segmentation maps. It presents SAGANet, a multimodal generative model that uses visual segmentation masks, video, and textual cues to produce controllable audio for musical instruments, offering fine‑grained, visually localized control. The authors also release the Segmented Music Solos dataset of instrument performance videos with segmentation information to support this task and demonstrate that SAGANet outperforms current state‑of‑the‑art methods in controllable, high‑fidelity Foley synthesis.
arXiv:2608.29073v1 Announce Type: new Abstract: Turn-by-turn (TBT) navigation systems are integral to modern driving experiences, providing real-time audio instructions to guide drivers safely to des...
arXiv:2610.00691v1 Announce Type: new Abstract: Recent joint audio-video generative models can synthesize realistic videos with synchronized sound, but typically generate audio as a single mixed trac...
arXiv:2607. 11364v1 Announce Type: cross Abstract: Generating immersive, synchronized and cinematic audio for long-form textual narratives remains a significant challenge in multi-modal AI.
Recent advances in generative video modeling have enabled diverse generation, reference-based synthesis, extension, and editing, but existing approaches often rely on fragmented task-specific models. A general model must distinguish heterogeneous target, source, and reference signals to determine what to generate, preserve, or use as guidance, while reducing interference among tasks.
arXiv:2609.38444v1 Announce Type: new Abstract: Text- and image-conditioned world generators can create visually rich 3D environments, yet these worlds often remain silent or rely on soundtracks synt...
Recent Large Audio Language Models (LALMs) have achieved remarkable progress in audio perceptual tasks across individual acoustic layers, including speech, sound, and music. However, existing benchmarks predominantly evaluate these layers in isolation, overlooking the complex contextual relationships that arise when multiple acoustic sources co-occur in real-world auditory scenes.
arXiv:2603. 07584v2 Announce Type: replace-cross Abstract: Computational engine sound modeling is central to the automotive audio industry, particularly for active sound design applications and virtual prototyping.
Music visualization offers a powerful way to enhance listeners' understanding and experience of music by translating auditory signals into visual forms. However, most existing approaches either rely heavily on lyrics or generate flat, non-immersive videos similar to conventional music videos, which limits their ability to convey the emotional dynamics of music and provide an immersive listening experience.