arXiv Machine Learning

MIME: Multimodal Interactive Motion Encoder

arXiv:2607. 22702v1 Announce Type: cross Abstract: Text-motion representation learning has advanced rapidly, with growing interest in multi person interactions for animation, AR/VR, and embodied AI.

arXiv Computer Vision
6d ago

Timo: $\textbf{T}$aming Mult$\textbf{i}$modal Diffusion Transformer for Human $\textbf{Mo}$tion Generation

The paper introduces Timo, a kinematics-aware multimodal diffusion transformer designed for human motion generation. Timo employs fully shared multimodal attention, flow matching, and geometric/rotational-kinematics supervision to better coordinate articulated motion, and uses a two-stage curriculum to align motion with text captions. The authors also present a new benchmark of 40,025 clips from six datasets, showing that Timo outperforms state‑of‑the‑art methods, achieving a 40.8% relative improvement over Kimodo on average.

By Zhao Wang, Jiangtao Hu, Jack Yu, Tao Yu
arXiv AI
Sep 15

Open-UniMo: Towards Unified Motion-Language Understanding and Generation in the Open World

arXiv:2609.14615v1 Announce Type: cross Abstract: Unified motion generation and understanding is crucial for embodied AI systems that can both synthesize and interpret human actions in open-world env...

By Guocun Wang, Kenkun Liu, Guorui Song, Jing Lin, Zhe Huang, Luyuan Zhang, Dake Zhong, Choo Sin Wai, Xiaoguang Han, Haoqian Wang
arXiv Computer Vision
Aug 26

SeMoCo: A Semantic-First Motion Codec for Motion Language Modeling

arXiv:2608.24334v1 Announce Type: new Abstract: Discrete motion representations have substantially advanced autoregressive text-to-motion generation. However, most motion tokenizers are optimized for...

By Tianlv Huang, Hetian Guo, Ziyi Cai, Song Wang, Yanping Zhang, Zipei Fan, Xuan Song, Guangming Wu, Xin Zheng
arXiv Computer Vision
Sep 14

Uni-HOI:A Unified framework for Learning the Joint distribution of Text and Human-Object Interaction

Uni-HOI is a unified framework that learns the joint distribution among text, human motion, and object motion for 4D human‑object interaction (HOI). It uses large language models and two motion‑specific VQ‑VAEs to convert heterogeneous motion data into token sequences, enabling seamless integration of all three modalities. A two‑stage training strategy first captures correlations on a large‑scale HOI dataset and then fine‑tunes for specific tasks, achieving strong performance on text‑driven HOI generation, object‑motion‑driven human motion generation, and human‑motion‑driven object motion prediction.

By Mengfei Zhang, Jinlu Zhang, Zhigang Tu
arXiv Computer Vision
Aug 28

Residual Flow Matching with Dynamic Cross-Interaction for 3D Multi-Person Motion Prediction

The paper introduces a Prior‑Guided Residual Flow Matching framework for 3D multi‑person motion prediction. It uses a Deterministic Coarse Prior to anchor kinematics and a Dynamic Cross‑Interaction mechanism to synchronize inter‑agent message passing during integration, thereby improving structural consistency and social context extraction. A decoupled joint‑motion architecture with bidirectional fusion further preserves fine‑grained kinematic coherence, achieving state‑of‑the‑art accuracy on several datasets.

By Wei Wei, Yinyuan Zhao, Ruixuan Yu
arXiv Computer Vision
Sep 25

SALI: Shot-Aware Late Interaction for Cross-Shot Relation Matching in Text-to-Video Retrieval using Film-Grammar Knowledge

SALI (Shot-Aware Late Interaction) is a new method for text-to-video retrieval that addresses the loss of relational information when videos are represented by a single embedding. It extracts the subject and object from a query sentence and matches their text embeddings against each visual shot embedding of a video clip using a greedy max or optimal transport operator, with a film-grammar penalty added during fine‑tuning. Built on CLIP4Clip‑meanP, SALI maintains overall recall on Condensed Movies and ActivityNet while significantly improving R@1 for multi‑shot relation queries, achieving the largest gains among compared methods and modestly improving performance on MSR‑VTT.

By Toya Oyama, Rainer Lienhart, Shin'ichi Satoh
arXiv AI
Jun 12

HYDRA-X: Native Unified Multimodal Models with Holistic Visual Tokenizers

arXiv:2606. 13289v1 Announce Type: cross Abstract: Holistic visual tokenizers are fundamental to unified multimodal models (UMMs) as they map diverse visual inputs into a unified representation space.

By Guozhen Zhang, Xuerui Qiu, Yutao Cui, Tianhui Song, Changlin Li, Junzhe Li, Tao Huang, Xiao Zhang, Yang Li, Jianbing Wu, Miles Yang, Zhao Zhong, Liefeng Bo, Limin Wang
arXiv AI
Aug 24

Identity-Aware Human-Object Interaction Motion Captioning

The paper introduces the Identity-Aware Human-Object Interaction Motion Captioning task, which requires captions to include both the subject’s identity and the interaction motion, e.g., "Sub_ID lifts the chair" instead of a generic description. It proposes ID‑HOINet, a model that learns from multi‑view videos using a Multi‑View Identity‑Motion Learning Module and a Two‑Stage Caption Rewriting Strategy to generate identity‑aware captions. Experiments show that ID‑HOINet achieves state‑of‑the‑art performance on the BEHAVE and InterCap datasets.

By Yiming Wang, Yonghao Dang, Huilai Li, Jiawei Tu, Jianqin Yin
arXiv Computer Vision
Sep 1

CL4D: Contrastive Language-4D Pretraining for Vision-Language Reasoning in Dynamic Scenes

arXiv:2608.18734v2 Announce Type: replace Abstract: 4D understanding and reasoning is a fundamental capability for embodied AI agents operating in dynamic physical environments. However, existing vis...

By Kumal Hewagamage, Isuranga Senavirathne, Sasika Amarasinghe, Hasitha Gallella, Dulanga Weerakoon, Vigneshwaran Subbaraju, Ranga Rodrigo