STyMo is a few‑shot motion style transfer method that learns from only seconds of paired data and trains in one to two minutes. It decomposes style into a static posture component and a temporal dynamics component, allowing runtime adjustment of posture intensity, temporal exaggeration, and per‑body‑region style. The approach includes a stylizability gate to avoid artifacts on out‑of‑distribution motions and supports an iterative authoring workflow, with results shown across a range of motion styles and a released dataset for future research.
By Jose Luis Ponton, Alexander Winkler, Ladislav Kavan, Yuting Ye, Petr Kadlecek
arXiv:2608.23279v1 Announce Type: new
Abstract: Text-driven human motion synthesis has made substantial development with two core modules of motion representation and generative architecture. For rep...
By Chengqun Yang, Liang Xu, Yanping Li, Fulong Liu, Jingnan Gao, Weili Zeng, Yichao Yan
arXiv:2609.08032v1 Announce Type: cross
Abstract: We introduce FlexMoGen, a novel framework for flexible human motion synthesis conditioned on both natural language descriptions and motion style refe...
By Kai Weixian Lan, Bodie Criswell, Briana Fedkiw, Zhan Zhang, Joseph Teran, Daniel Holden
The paper introduces a unified conditional-flow framework that integrates text-driven motion generation, semantic editing, and intra-structural retargeting into a single rectified-flow model. By treating editing as a change in semantic condition and retargeting as a change in skeletal condition, the approach eliminates fragmented pipelines and allows a single model to perform generation, zero‑shot editing, and zero‑shot retargeting on articulated 3D motion data. Experiments on SnapMoGen and a Mixamo subset demonstrate that the model can handle all three tasks without task‑specific fine‑tuning, preserving both motion semantics and skeletal structure.
By Junlin Li, Xinhao Song, Siqi Wang, Haibin Huang, Yili Zhao
arXiv:2603.08590v4 Announce Type: replace
Abstract: Text-to-motion generation has advanced with larger corpora and stronger generators, yet many models still rely on holistic frame- or clip-level lat...
By Zeyu Ling, Qing Shuai, Teng Zhang, Shiyang Li, Bo Han, Changqing Zou
The paper introduces Motion Style Slider, a framework that enables continuous, endpoint‑supervised control of style intensity in human motion diffusion. By constructing a style direction in a learned motion‑style embedding space and conditioning diffusion generation with a scalar intensity, the method achieves smooth, monotonic style scaling without needing intermediate‑intensity ground truth. The approach is compatible with pretrained diffusion backbones, supports heterogeneous style datasets, and is evaluated on controllability, interpolation/extrapolation, content preservation, and motion realism.
By Chen-Chieh Liao, Yichen Peng, Yiyi Cai, Y\^ui Ono, Hiroki Hanaoka, Erwin Wu, Hideki Koike, Shuichi Kurabayashi
arXiv:2609.23817v1 Announce Type: new
Abstract: We present VISTA, a two-stage framework for generating stylized 3D human motion by fusing structural content from text prompts with expressive style fr...
By Monseej Purkayastha, Anindita Ghosh, Philipp Slusallek
arXiv:2509. 15443v2 Announce Type: replace-cross Abstract: Human-to-humanoid imitation learning presents a promising pathway to address the severe data scarcity bottleneck in robotics by utilizing abundant, large-scale human motion collections.
By Xingyu Chen, Hanyu Wu, Sikai Wu, Mingliang Zhou, Diyun Xiang, Haodong Zhang, Yangchen Zhou, Yukang Gao, Yi Gu, Renjing Xu
arXiv:2607. 08741v1 Announce Type: cross Abstract: Generating realistic 3D human motions in real-time within interactive applications is key for animation, simulation, and humanoid robotics.
By Kaifeng Zhao, Mathis Petrovich, Haotian Zhang, Tingwu Wang, Siyu Tang, Davis Rempe
arXiv:2608.20699v1 Announce Type: new
Abstract: Animating articulated 3D meshes via text requires satisfying strict kinematic constraints, modeling causal interactions between parts, and achieving in...
By Chunyu Zou, Peng Dai, Yi-Hua Huang, Ze Yuan, Jingwei Huang, Yeming Yao, Xiaojuan Qi
arXiv:2606. 31981v1 Announce Type: cross Abstract: Creating photorealistic, animatable 3D human avatars from monocular images still largely depends on Linear Blend Skinning (LBS) and parametric body models, which constrain expressivity and often introduce artifacts due to imperfect fitting.
By Peng Li, Rawal Khirodkar, Junxuan Li, Yuan Dong, Chen Cao, Yuan Liu, Wenhan Luo, Yike Guo, Shunsuke Saito
Instruction-driven editing of 3D human motion requires precise spatiotemporal localization, rich semantic grounding, and strict preservation of unmodified content. Existing methods either resort to training-free adaptation of generative models or rely solely on triplet supervision; however, adaptation often yields suboptimal control, and manually curated triplet datasets remain severely limited in scale and semantic diversity.