The paper compares diffusion and rectified flow objectives within the MotionGPT3 framework for text-driven motion generation. Experiments on HumanML3D show that rectified flow converges faster, achieves strong test performance earlier, and matches or exceeds diffusion quality while requiring fewer sampling steps. The study isolates the generative objective’s impact, demonstrating that rectified flow’s benefits transfer to continuous-latent motion generation.
By Jaymin Bhan, JiHong Jeon, SangYeop Jeong
arXiv:2607. 29180v1 Announce Type: cross Abstract: Text-to-motion generation must produce motions that are semantically correct, temporally coherent, and physically plausible.
By Yifei Zhu, Mingyi Shi, Yangyang Cai, Miao Cheng, Yoshifumi Kitamura, Taku Komura
The paper introduces Timo, a kinematics-aware multimodal diffusion transformer designed for human motion generation. Timo employs fully shared multimodal attention, flow matching, and geometric/rotational-kinematics supervision to better coordinate articulated motion, and uses a two-stage curriculum to align motion with text captions. The authors also present a new benchmark of 40,025 clips from six datasets, showing that Timo outperforms state‑of‑the‑art methods, achieving a 40.8% relative improvement over Kimodo on average.
By Zhao Wang, Jiangtao Hu, Jack Yu, Tao Yu
arXiv:2603. 22282v2 Announce Type: replace-cross Abstract: We present UniMotion, to our knowledge the first unified framework for simultaneous understanding and generation of human motion, natural language, and RGB images within a single architecture.
By Ziyi Wang, Xinshun Wang, Shuang Chen, Yang Cong, Mengyuan Liu
arXiv:2609.14615v1 Announce Type: cross
Abstract: Unified motion generation and understanding is crucial for embodied AI systems that can both synthesize and interpret human actions in open-world env...
By Guocun Wang, Kenkun Liu, Guorui Song, Jing Lin, Zhe Huang, Luyuan Zhang, Dake Zhong, Choo Sin Wai, Xiaoguang Han, Haoqian Wang
Diffusion-based text-to-motion models synthesize realistic human motions but often exhibit semantic drift from the input text. Motion is inherently temporal, especially in compositional and long-duration sequences that require semantic consistency across multiple action segments and smooth kinematic transitions throughout the trajectory.