arXiv AI By Akira Takahashi, Chihiro Nagashima, Zhi Zhong, Shusuke Takahashi, Yuki Mitsufuji

Passing: An Endless Journey through Reconstructed Spacetime with AI-Generated Sound

Read the original on arXiv AI →

Passing is an interactive audiovisual installation that transforms a single continuous monorail-window recording into an endless journey by reconstructing it as a spatiotemporal volume and resampling its spatial and temporal structure along nonlinear trajectories. A camera-based viewer‑presence detection system influences transitions among rendered video sequences, and the resulting video stream is fed into SpecMaskFoley, a real‑time video‑to‑audio synthesis model that generates a synchronized soundscape. The work distributes creative agency among the artist, the AI model, and the audience, exploring how authorship and listening can be negotiated among human intention, machine inference, and audience interpretation.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Jun 2

Do Joint Audio-Video Generation Models Understand Physics?

arXiv:2605. 07061v2 Announce Type: replace-cross Abstract: Joint audio-video generation models are rapidly approaching professional production quality, raising a central question: do they understand audio-visual physics, or merely generate plausible sounds and frames that violate real-world consistency?

By Zijun Cui, Xiulong Liu, Hao Fang, Mingwei Xu, Jiageng Liu, Zexin Xu, Weiguo Pian, Shijian Deng, Feiyu Du, Chenming Ge, Yapeng Tian
arXiv Computer Vision
Sep 21

Personalizing Causal Audio-Driven Facial Motion via Dynamic Multi-modal Retrieval

The paper introduces an end‑to‑end framework for personalized audio‑driven facial motion that operates in real time without look‑ahead. It combines a causal multi‑resolution motion tokenizer, which captures both global temporal context and fine articulatory details, with a multi‑modal style retriever that pulls stylistic priors from arbitrary reference footage using ongoing audio and motion queries. This approach allows high‑fidelity, identity‑consistent animation from just a few casually recorded clips, outperforming existing methods in lip‑sync accuracy, identity consistency, and perceived realism while maintaining real‑time streaming constraints.

By Xuangeng Chu, Yu Han, Wei Mao, Shih-En Wei