arXiv:2506. 01982v5 Announce Type: replace-cross Abstract: This study investigates emotional expression and perception in music performance using computational and neurophysiological methods.
By Vassilis Lyberatos, Spyridon Kantarelis, Ioanna Zioga, Christina Anagnostopoulou, Giorgos Stamou, Anastasia Georgaki
Music visualization offers a powerful way to enhance listeners' understanding and experience of music by translating auditory signals into visual forms. However, most existing approaches either rely heavily on lyrics or generate flat, non-immersive videos similar to conventional music videos, which limits their ability to convey the emotional dynamics of music and provide an immersive listening experience.
Audio-driven talking head synthesis has achieved impressive progress in lip synchronization and visual quality, yet generating expressive emotional avatars with controllable intensity remains challenging, especially under real-time constraints. In this paper, we present GaussianEmoTalker, an audio-driven framework for real-time emotional talking head synthesis based on 3D Gaussian Splatting.
Passing is an interactive audiovisual installation that transforms a single continuous monorail-window recording into an endless journey by reconstructing it as a spatiotemporal volume and resampling its spatial and temporal structure along nonlinear trajectories. A camera-based viewer‑presence detection system influences transitions among rendered video sequences, and the resulting video stream is fed into SpecMaskFoley, a real‑time video‑to‑audio synthesis model that generates a synchronized soundscape. The work distributes creative agency among the artist, the AI model, and the audience, exploring how authorship and listening can be negotiated among human intention, machine inference, and audience interpretation.
By Akira Takahashi, Chihiro Nagashima, Zhi Zhong, Shusuke Takahashi, Yuki Mitsufuji
EmoTra‑TTS introduces a method for smooth intra‑utterance emotion transitions in speech synthesis. It uses a multi‑pass flow blending pipeline, dual‑stage VAD conditioning, and direction‑magnitude decoupled injection to generate frame‑aligned emotional prosody. The system adds only 0.43% more parameters, incurs no latency, and outperforms four state‑of‑the‑art baselines and two commercial systems in emotion transition quality and overall preference tests.
By Tianchi Liu, Zeyang Song, Tianrui Wang, Zhipeng Li, Chenglin Xu, Yiwen Guo
arXiv:2608. 08349v1 Announce Type: cross Abstract: Audio dramas weave dialogue, sound effects, and music into immersive stories.
By Karim Benharrak, Oriol Nieto, Bryan Wang, Zeyu Jin, Amy Pavel