arXiv AI

Music Interpretation and Emotion Perception: A Computational and Neurophysiological Investigation

arXiv:2506. 01982v5 Announce Type: replace-cross Abstract: This study investigates emotional expression and perception in music performance using computational and neurophysiological methods.

arXiv AI
Jul 8

From Textural Counterpoint to Feature Encoding: A Multi-Dimensional Machine Representation Study of Haydn's "The Lark" Integrating Electroacoustic Analysis

arXiv:2607. 05902v1 Announce Type: cross Abstract: Chamber music, as a highly precise multi-part interactive system, contains a logic of "role assignment and dynamic interaction" that provides an extremely valuable blueprint for exploring human-computer collaborative composition paradigms.

By Yakun Liu, Zhiyu Jin, Hai Luan, Dong Liu, Xiaonan Li
arXiv AI
Aug 17

Musical Agent Systems: MACAT and MACataRT

arXiv:2502. 00023v2 Announce Type: replace-cross Abstract: Our research explores the development and application of musical agents, human-in-the-loop generative AI systems designed to support music performance and improvisation within co-creative spaces.

By Keon Ju M. Lee, Philippe Pasquier
arXiv Computer Vision
Sep 7

InterSing: Explicit Interaction Dynamics for 3D Duet Singing Animation and Beyond

InterSing is a framework that generates realistic 3D head animations for duet singing by modeling the sparse, rhythm‑dependent interactions between performers. It introduces interaction logits—a weakly supervised, interpretable latent representation of cross‑performer engagement—and uses them to condition an interaction‑aware diffusion model driven by audio and interaction dynamics. The approach enables unified multi‑mode generation, producing coordinated behavior, independent motion, and smooth transitions, and it generalizes to multi‑singer performances with intuitive control over engagement.

By Yihan Zhou, Zikai Huang, Yuyang Yu, Xuemiao Xu, Cheng Xu, Shengfeng He
arXiv Machine Learning
Sep 16

EMODY Flow: Emotion-Aware Audio-Driven Full-Body Motion Generation

EMODY Flow is a lightweight flow‑matching framework that generates synchronized full‑body motion and facial expressions conditioned on speech and emotion. It attaches to a frozen Qwen‑3 Omni model, reusing its audio codecs to drive two parallel DiT generators for SMPL‑X body pose and FLAME facial expressions. An auxiliary emotion classifier at training time restores emotion sensitivity, enabling EMODY Flow to achieve state‑of‑the‑art gesture quality on BEAT2 and zero‑shot facial animation on TFHP, with significant improvements in FGD, Beat Correlation, and Diversity metrics.

By Harsh Kumar Agarwal, Xavier Alameda-Pineda, Olivier Perrotin