The paper explores person identification using millimeter‑wave point clouds beyond traditional gait analysis, focusing on seven activities of daily living (ADLs). It introduces the mm‑ADL dataset of 11 subjects and proposes an activity‑conditioned framework that routes each clip to an activity‑specific identity expert via a supervised mixture of experts. Experiments show that hard routing improves closed‑set ID accuracy from 62.1% to 68.0% and significantly boosts re‑identification metrics, demonstrating the benefit of activity context under controlled indoor conditions.
The paper introduces the Identity-Aware Human-Object Interaction Motion Captioning task, which requires captions to include both the subject’s identity and the interaction motion, e.g., "Sub_ID lifts the chair" instead of a generic description. It proposes ID‑HOINet, a model that learns from multi‑view videos using a Multi‑View Identity‑Motion Learning Module and a Two‑Stage Caption Rewriting Strategy to generate identity‑aware captions. Experiments show that ID‑HOINet achieves state‑of‑the‑art performance on the BEHAVE and InterCap datasets.
By Yiming Wang, Yonghao Dang, Huilai Li, Jiawei Tu, Jianqin Yin
arXiv:2608.21093v1 Announce Type: new
Abstract: Stochastic human motion prediction aims to forecast future motion distributions. Although recent studies have achieved strong performance in terms of a...
By Yue Ma, Frederick W. B. Li, Xiaohui Liang
arXiv:2604. 18064v2 Announce Type: replace Abstract: Human motion world models should capture motion's intentionality by being executable: adaptable to different actions and capable of assessing motion quality.
By Rimvydas Rubavicius, Manisha Dubey, N. Siddharth, Subramanian Ramamoorthy
arXiv:2607. 29181v1 Announce Type: cross Abstract: Agentic assistants capable of proactive, personalized interactions require structured models of user intent and workflow.
By Andy J. Phu, James Mooney, Karin de Langis, Khanh Chi Le, Dongyeop Kang
arXiv:2503. 15225v3 Announce Type: replace-cross Abstract: The deployment of autonomous virtual avatars (in extended reality) and robots in human group activities---such as rehabilitation therapy, sports, and manufacturing---is expected to increase as these technologies become more pervasive.
By Angelo Di Porzio, Marco Coraggio
arXiv:2609.18432v1 Announce Type: new
Abstract: Extensive occlusions in real-world scenarios pose challenges to gait recognition due to missing and noisy information, as well as body misalignment in...
By Panjian Huang, Yunjie Peng, Saihui Hou, Chunshui Cao, Xu Liu, Zhiqiang He, Yongzhen Huang
arXiv:2609.14118v1 Announce Type: cross
Abstract: Online understanding of who is talking to the camera wearer is a key capability for egocentric social interaction. However, existing talk-to-me (TTM)...
By Feiyu Du, Xi He, Jia Li, Yapeng Tian, Weili Wu
Chehre is an emoji‑prompted video dataset designed to study perceptual flexibility in video language models. It contains 2,111 videos of 203 participants expressing 40 facial emojis, with each video annotated by about 30 perceivers, yielding 1,242 annotators in total. The dataset introduces a new task—distributional expression recognition—that evaluates a model’s ability to reproduce the variation seen in human annotations, and shows that persona prompting can shift model perception to better match human variability.
By Bita Azari, Zoe Stanley, Avneet Batra, Poorvi Bhatia, Hali Kil, Manolis Savva, Angelica Lim
The paper introduces a Prior‑Guided Residual Flow Matching framework for 3D multi‑person motion prediction. It uses a Deterministic Coarse Prior to anchor kinematics and a Dynamic Cross‑Interaction mechanism to synchronize inter‑agent message passing during integration, thereby improving structural consistency and social context extraction. A decoupled joint‑motion architecture with bidirectional fusion further preserves fine‑grained kinematic coherence, achieving state‑of‑the‑art accuracy on several datasets.
By Wei Wei, Yinyuan Zhao, Ruixuan Yu
arXiv:2608. 05115v1 Announce Type: cross Abstract: Can computer vision help make classrooms safer?
By Paritosh Parmar, Landy Lan, Hong Yang, Chen Yi, Chiat Pin Tay
arXiv:2606. 28769v1 Announce Type: new Abstract: Emotional body motion expressions are an essential element of non-verbal communication.
By Huakun Liu, Miao Cheng, Xin Wei, Felix Dollack, Victor Schneider, Hideaki Uchiyama, Chia-huei Tseng, Yoshifumi Kitamura, Monica Perusquia-Hernandez