arXiv AI

EMORSION: Examining the Impact of Audio Parameters on Emotional Responses and Immersion in Film

arXiv:2606. 18266v1 Announce Type: cross Abstract: EMORSION is an exploratory proof-of-concept study examining how film audio design shapes audience emotion and immersion in acinema setting.

Hugging Face Trending Papers
Jul 15

Bring Music The Horizon: Music-Driven 360$^\circ$ Video Generation

Music visualization offers a powerful way to enhance listeners' understanding and experience of music by translating auditory signals into visual forms. However, most existing approaches either rely heavily on lyrics or generate flat, non-immersive videos similar to conventional music videos, which limits their ability to convey the emotional dynamics of music and provide an immersive listening experience.

Hugging Face Trending Papers
Jul 1

GaussianEmoTalker: Real-Time Emotional Talking Head Synthesis with Audio-Driven and Blendshape-Based 3D Gaussian Splatting

Audio-driven talking head synthesis has achieved impressive progress in lip synchronization and visual quality, yet generating expressive emotional avatars with controllable intensity remains challenging, especially under real-time constraints. In this paper, we present GaussianEmoTalker, an audio-driven framework for real-time emotional talking head synthesis based on 3D Gaussian Splatting.

arXiv AI
Sep 24

Passing: An Endless Journey through Reconstructed Spacetime with AI-Generated Sound

Passing is an interactive audiovisual installation that transforms a single continuous monorail-window recording into an endless journey by reconstructing it as a spatiotemporal volume and resampling its spatial and temporal structure along nonlinear trajectories. A camera-based viewer‑presence detection system influences transitions among rendered video sequences, and the resulting video stream is fed into SpecMaskFoley, a real‑time video‑to‑audio synthesis model that generates a synchronized soundscape. The work distributes creative agency among the artist, the AI model, and the audience, exploring how authorship and listening can be negotiated among human intention, machine inference, and audience interpretation.

By Akira Takahashi, Chihiro Nagashima, Zhi Zhong, Shusuke Takahashi, Yuki Mitsufuji
arXiv AI
Aug 26

EmoTra-TTS: Smooth Intra-Utterance Emotion Transitions for Speech Synthesis

EmoTra‑TTS introduces a method for smooth intra‑utterance emotion transitions in speech synthesis. It uses a multi‑pass flow blending pipeline, dual‑stage VAD conditioning, and direction‑magnitude decoupled injection to generate frame‑aligned emotional prosody. The system adds only 0.43% more parameters, incurs no latency, and outperforms four state‑of‑the‑art baselines and two commercial systems in emotion transition quality and overall preference tests.

By Tianchi Liu, Zeyang Song, Tianrui Wang, Zhipeng Li, Chenglin Xu, Yiwen Guo
arXiv AI
Jun 2

Do Joint Audio-Video Generation Models Understand Physics?

arXiv:2605. 07061v2 Announce Type: replace-cross Abstract: Joint audio-video generation models are rapidly approaching professional production quality, raising a central question: do they understand audio-visual physics, or merely generate plausible sounds and frames that violate real-world consistency?

By Zijun Cui, Xiulong Liu, Hao Fang, Mingwei Xu, Jiageng Liu, Zexin Xu, Weiguo Pian, Shijian Deng, Feiyu Du, Chenming Ge, Yapeng Tian