arXiv:2606. 22726v2 Announce Type: replace Abstract: Choreographic motion generation poses unique challenges for AI, demanding precise semantic control over complex, temporally structured, and expressive full-body dynamics.
By Seong Jong Yoo, Siyuan Peng, Felix Gu, Stratis Aloimonos, Cornelia Ferm\"uller
STyMo is a few‑shot motion style transfer method that learns from only seconds of paired data and trains in one to two minutes. It decomposes style into a static posture component and a temporal dynamics component, allowing runtime adjustment of posture intensity, temporal exaggeration, and per‑body‑region style. The approach includes a stylizability gate to avoid artifacts on out‑of‑distribution motions and supports an iterative authoring workflow, with results shown across a range of motion styles and a released dataset for future research.
By Jose Luis Ponton, Alexander Winkler, Ladislav Kavan, Yuting Ye, Petr Kadlecek
arXiv:2606. 01014v1 Announce Type: cross Abstract: We address text-based 3D human motion editing, where the goal is to preserve the style and structure of a source motion while applying edits described in natural language.
By Gyojin Han, Junmo Kim
arXiv:2601. 08828v2 Announce Type: replace-cross Abstract: Despite the rapid progress of video generation models, the role of data in influencing motion is poorly understood.
By Xindi Wu, Despoina Paschalidou, Jun Gao, Antonio Torralba, Laura Leal-Taix\'e, Olga Russakovsky, Sanja Fidler, Jonathan Lorraine
Instruction-driven editing of 3D human motion requires precise spatiotemporal localization, rich semantic grounding, and strict preservation of unmodified content. Existing methods either resort to training-free adaptation of generative models or rely solely on triplet supervision; however, adaptation often yields suboptimal control, and manually curated triplet datasets remain severely limited in scale and semantic diversity.
The paper compares diffusion and rectified flow objectives within the MotionGPT3 framework for text-driven motion generation. Experiments on HumanML3D show that rectified flow converges faster, achieves strong test performance earlier, and matches or exceeds diffusion quality while requiring fewer sampling steps. The study isolates the generative objective’s impact, demonstrating that rectified flow’s benefits transfer to continuous-latent motion generation.
By Jaymin Bhan, JiHong Jeon, SangYeop Jeong
arXiv:2605. 29488v2 Announce Type: replace-cross Abstract: Conditional human motion generation remains a fundamental challenge in computer vision and robotics.
By Yiheng Li, Zhuo Li, Ruibing Hou, Yingjie Chen, Hong Chang, Hao Liu, Shiguang Shan
arXiv:2608.24334v1 Announce Type: new
Abstract: Discrete motion representations have substantially advanced autoregressive text-to-motion generation. However, most motion tokenizers are optimized for...
By Tianlv Huang, Hetian Guo, Ziyi Cai, Song Wang, Yanping Zhang, Zipei Fan, Xuan Song, Guangming Wu, Xin Zheng
Music-driven dance video generation aims to synthesize expressive human motion that is temporally aligned with music while maintaining high visual fidelity. Despite recent progress, existing methods still face two key limitations: the lack of large-scale, high-quality dance video datasets, and the absence of principled frameworks for integrating music as a complementary conditioning signal into Video Generation Foundation Models.
arXiv:2608.23279v1 Announce Type: new
Abstract: Text-driven human motion synthesis has made substantial development with two core modules of motion representation and generative architecture. For rep...
By Chengqun Yang, Liang Xu, Yanping Li, Fulong Liu, Jingnan Gao, Weili Zeng, Yichao Yan
Recent advances in generative video modeling have enabled diverse generation, reference-based synthesis, extension, and editing, but existing approaches often rely on fragmented task-specific models. A general model must distinguish heterogeneous target, source, and reference signals to determine what to generate, preserve, or use as guidance, while reducing interference among tasks.
arXiv:2606. 30266v1 Announce Type: cross Abstract: Motion-language agents must possess the bidirectional capability to both understand human movement (motion-to-text, M2T) and generate it from natural language (text-to-motion, T2M).
By Bertram Taetz, Hugo Albuquerque Cosme da Silva, Gabriele Bleser-Taetz