arXiv Computer Vision

MorphoStyle: Motion Style Transfer with Morphology Control

MorphoStyle is a new framework for shape‑aware motion style transfer that uses a shape‑conditioned FSQ‑VAE. It disentangles style from content through a contrastive style encoder, a text‑guided style‑routing mechanism, and a manifold‑preserving style modulator. Experiments on benchmark datasets show that MorphoStyle outperforms existing baselines in both shape control and motion style transfer.

arXiv Computer Vision
Sep 7

STyMo: Fast and Controllable Few-Shot Motion Style Transfer

STyMo is a few‑shot motion style transfer method that learns from only seconds of paired data and trains in one to two minutes. It decomposes style into a static posture component and a temporal dynamics component, allowing runtime adjustment of posture intensity, temporal exaggeration, and per‑body‑region style. The approach includes a stylizability gate to avoid artifacts on out‑of‑distribution motions and supports an iterative authoring workflow, with results shown across a range of motion styles and a released dataset for future research.

By Jose Luis Ponton, Alexander Winkler, Ladislav Kavan, Yuting Ye, Petr Kadlecek
arXiv AI
Aug 28

A Unified Conditional Flow for Motion Generation, Editing, and Intra-Structural Retargeting

The paper introduces a unified conditional-flow framework that integrates text-driven motion generation, semantic editing, and intra-structural retargeting into a single rectified-flow model. By treating editing as a change in semantic condition and retargeting as a change in skeletal condition, the approach eliminates fragmented pipelines and allows a single model to perform generation, zero‑shot editing, and zero‑shot retargeting on articulated 3D motion data. Experiments on SnapMoGen and a Mixamo subset demonstrate that the model can handle all three tasks without task‑specific fine‑tuning, preserving both motion semantics and skeletal structure.

By Junlin Li, Xinhao Song, Siqi Wang, Haibin Huang, Yili Zhao
arXiv Computer Vision
4d ago

Motion Style Slider: Endpoint-Supervised Continuous Style Control for Human Motion Diffusion

The paper introduces Motion Style Slider, a framework that enables continuous, endpoint‑supervised control of style intensity in human motion diffusion. By constructing a style direction in a learned motion‑style embedding space and conditioning diffusion generation with a scalar intensity, the method achieves smooth, monotonic style scaling without needing intermediate‑intensity ground truth. The approach is compatible with pretrained diffusion backbones, supports heterogeneous style datasets, and is evaluated on controllability, interpolation/extrapolation, content preservation, and motion realism.

By Chen-Chieh Liao, Yichen Peng, Yiyi Cai, Y\^ui Ono, Hiroki Hanaoka, Erwin Wu, Hideki Koike, Shuichi Kurabayashi
arXiv AI
Jul 1

A Scalable Whole-body Motion Transfer via Implicit Kinodynamic Motion Retargeting

arXiv:2509. 15443v2 Announce Type: replace-cross Abstract: Human-to-humanoid imitation learning presents a promising pathway to address the severe data scarcity bottleneck in robotics by utilizing abundant, large-scale human motion collections.

By Xingyu Chen, Hanyu Wu, Sikai Wu, Mingliang Zhou, Diyun Xiang, Haodong Zhang, Yangchen Zhou, Yukang Gao, Yi Gu, Renjing Xu
arXiv AI
Jul 1

LUNA: Learning Universal 3D Human Animation Beyond Skinning

arXiv:2606. 31981v1 Announce Type: cross Abstract: Creating photorealistic, animatable 3D human avatars from monocular images still largely depends on Linear Blend Skinning (LBS) and parametric body models, which constrain expressivity and often introduce artifacts due to imperfect fitting.

By Peng Li, Rawal Khirodkar, Junxuan Li, Yuan Dong, Chen Cao, Yuan Liu, Wenhan Luo, Yike Guo, Shunsuke Saito
Hugging Face Trending Papers
Aug 10

UniMoFlow: Grounding Instruction-Driven 3D Human Motion Editing in Generation

Instruction-driven editing of 3D human motion requires precise spatiotemporal localization, rich semantic grounding, and strict preservation of unmodified content. Existing methods either resort to training-free adaptation of generative models or rely solely on triplet supervision; however, adaptation often yields suboptimal control, and manually curated triplet datasets remain severely limited in scale and semantic diversity.