arXiv:2608.23279v1 Announce Type: new
Abstract: Text-driven human motion synthesis has made substantial development with two core modules of motion representation and generative architecture. For rep...
By Chengqun Yang, Liang Xu, Yanping Li, Fulong Liu, Jingnan Gao, Weili Zeng, Yichao Yan
The paper introduces a promptable localized motion representation that generates persistent embeddings for user-specified regions in a video, without cropping or masking the input. By conditioning motion encoding directly on spatial masks while processing the full video, the method produces temporally consistent, region-addressable embeddings that capture local dynamics while preserving global context. These embeddings enable object-level motion transfer for dynamic scene composition and improve localized action classification in multi-actor videos, outperforming global representations that rely on cropping or post-hoc masking.
By Frank Fundel, Malek Ben Alaya, Thomas Ressler-Antal, Stefan Andreas Baumann, Bj\"orn Ommer
The paper introduces a latent dataset distillation framework for human motion prediction, addressing the limitations of traditional gradient matching by incorporating a learned motion prior. Motions are compressed using a residual‑quantized variational autoencoder, and distillation updates only a latent bank while keeping the decoder frozen, ensuring synthetic motions remain plausible. Experiments on Human3.6M, CMU, and 3DPW datasets demonstrate that this method outperforms direct gradient matching in most settings and yields more realistic synthetic motions.
By Ge Tian, Guang Li, Takahiro Ogawa, Miki Haseyama
arXiv:2607. 17342v1 Announce Type: cross Abstract: Understanding physical human-robot and human-human interactions is a challenging yet emerging topic in 3D vision.
By Yuhang Wen, Mengyuan Liu, Zixuan Tang, Junsong Yuan, Sirui Li, Beichen Ding
arXiv:2608. 19987v1 Announce Type: new Abstract: Skeleton-based Video Anomaly Detection (VAD) offers a robust, privacy-preserving solution for identifying abnormal behaviors.
By Jakub Micorek, Mateusz Kozi\'nski, Horst Possegger
arXiv:2607. 16322v1 Announce Type: cross Abstract: Micro-gesture recognition demands the detection of fleeting, spatially localized movements that are frequently overwhelmed by dominant static appearances and background noise.
By Taorui Wang, Wei Xia, Hui Ma, Zijia Song, Jiayu Zhang, Zeheng Wang, Yong Xu, Zitong Yu
The paper introduces Skeleton-Language feature Pooling Switching, a weakly‑supervised vision‑language pretraining strategy for skeleton‑based zero‑shot spatio‑temporal action localization. It replaces video‑level pooling with instance‑level feature computation during inference, enabling the model to estimate unseen actions without costly annotations. Additionally, Scene‑Mixed Discriminative Contrastive Learning is proposed to separate actions at the instance level within mixed scenes using a MIL framework, and experiments on four public datasets confirm the method’s effectiveness.
By Koshiro Nagano, Fumiaki Sato, Ryo Hachiuma, Kazuki Tsutsukawa, Taiki Sekii
arXiv:2607. 15400v1 Announce Type: cross Abstract: Falls among older adults are a major safety challenge, but continuous monitoring is difficult to sustain.
By Tasmiah Haque, Jacob Kosinski, Sumit Mohan, Srinjoy Das, Mohammad Abdullah Al-Mamun
arXiv:2607. 27581v1 Announce Type: new Abstract: Grounding human motion in language, and language in motion, is a central step toward physical AI systems that can understand, generate, and communicate human behavior.
By Zhankai Ye, Yukai Jin, Bingyang Wei, Bofan Li, Yusen Wu, Fangyi Li, Shangqian Gao, Xin Liu
arXiv:2609.37495v1 Announce Type: new
Abstract: Human motion generation plays an important role in applications such as character animation, virtual environments, and embodied interaction. While exis...
By Yun Chen, Munchurl Kim, Jeonghyeok Do
arXiv:2606. 29531v1 Announce Type: cross Abstract: We propose MotionAtlas, a system for detailed captioning of motion-centric videos, comprising (1) a dedicated human-annotated benchmark, (2) a scalable, high-quality pipeline to construct training samples, and (3) a family of powerful Video-MLLMs.
By Weisong Liu, Haochen Wang, Kuan Gao, Yuhao Wang, Yikang Zhou, Zhongwei Ren, Jacky Mai, Anna Wang, Yanwei Li, Jason Li, Zhaoxiang Zhang
arXiv:2603. 22282v2 Announce Type: replace-cross Abstract: We present UniMotion, to our knowledge the first unified framework for simultaneous understanding and generation of human motion, natural language, and RGB images within a single architecture.
By Ziyi Wang, Xinshun Wang, Shuang Chen, Yang Cong, Mengyuan Liu