MMGait is a large‑scale multi‑sensor benchmark that aligns visible, infrared, depth, LiDAR, and radar observations at the sequence level, enabling evaluation of single‑modal, cross‑modal, and multi‑modal gait recognition. The study shows that modality rankings shift with probe conditions, cross‑modal alignment remains challenging, and fusion can yield complementary gains. To address the scalability issue of training separate experts, the authors propose Omni‑Modal Gait Recognition and its implementation, OmniGait++, which unifies all recognition settings within a shared identity space using modality‑specific front ends, a shared encoder, and an anchor‑guided fusion module.
whyItMatters":"MMGait provides a common testbed for heterogeneous gait sensing and demonstrates that unified recognition across varying modality availability is feasible, offering a scalable alternative to task‑specific experts."
By Saihui Hou, Chenye Wang, Qingyuan Cai, Aoqi Li, Yongzhen Huang
Action Quality Assessment (AQA) aims to objectively evaluate performance quality from action videos. Most existing methods follow a ``one-by-one'' paradigm, training a separate model for each action type.
arXiv:2410. 19553v2 Announce Type: replace-cross Abstract: This paper explores the impact of occlusions in video action detection.
By Rajat Modi, Vibhav Vineet, Yogesh Singh Rawat
arXiv:2606. 31127v1 Announce Type: cross Abstract: To enable personalized, real-time coaching using Augmented Reality glasses or fixed camera setups in domains such as sports, cooking, or music, a system must understand not just what a person does, but how well they execute an activity.
By Bj\"orn Braun, Christian Holz
The paper introduces Composed Gait Retrieval (CoGR), a task that retrieves a target gait sequence using a reference sequence and a natural language modification query. To support this, the authors create the first gait-language datasets—Language‑Augmented CCPG and Language‑Augmented CASIA‑B—via an automated annotation pipeline powered by large vision‑language models. They propose ComposeGait, an identity‑anchored composition framework with a Part‑aware Identity Adapter that injects identity tokens into a shared Q‑Former, achieving state‑of‑the‑art retrieval performance on both benchmarks.
By Jingchen Fei, Zengbin Wang, Yukun Liu, Muyi Sun, Shibiao Xu, Man Zhang
arXiv:2609.18413v1 Announce Type: new
Abstract: What is a gait? Appearance-based gait networks consider a gait as the human shape and motion information from images. Model-based gait networks treat a...
By Panjian Huang, Saihui Hou, Chunshui Cao, Xu Liu, Yongzhen Huang
arXiv:2609.08038v2 Announce Type: cross
Abstract: Smart healthcare monitoring systems require precise action recognition to ensure well-being and timely intervention in critical situations such as fa...
By Diwas Lamsal, Pramod Wickramatilake, Jednipat Moonrinta, Mongkol Ekpanyapong, Matthew N. Dailey
arXiv:2609.10187v1 Announce Type: new
Abstract: In this work, we introduce language-aligned motion representations for domain-generalizable UPDRS-Gait severity estimation, aiming to learn semanticall...
By Soojie Kim, Muhammad Munsif, Minkyung Kim, Seungryul Baek
arXiv:2607. 17342v1 Announce Type: cross Abstract: Understanding physical human-robot and human-human interactions is a challenging yet emerging topic in 3D vision.
By Yuhang Wen, Mengyuan Liu, Zixuan Tang, Junsong Yuan, Sirui Li, Beichen Ding
arXiv:2608. 10932v1 Announce Type: cross Abstract: Understanding camera motion is fundamental to video perception, with applications in spatial intelligence and controllable video generation.
By Dazhao Du, Shiyan Du, Jian Liu, Yongjian Yu, Bohai Gu, Tao Han, Hualuo Liu, Eric Liu, Yujia Zhang, Xi Chen, Song Guo
arXiv:2606. 00054v1 Announce Type: cross Abstract: Recent progress in generalizable embodied control has been driven by large-scale pretraining of Vision-Language-Action (VLA) models.
By Zhiyuan Feng, Qixiu Li, Huizhi Liang, Rushuai Yang, Yichao Shen, Zhiying Du, Zhaowei Zhang, Yu Deng, Li Zhao, Hao Zhao, Zongqing Lu, Oier Mees, Marc Pollefeys, Jiaolong Yang, Baining Guo
Fine-grained understanding of operating room (OR) activity could enable workflow-aware assistance, yet remains difficult due to clutter, occlusions, and limited sensing. The prevailing approach to model this environment is scene graphs as an interpretable representation of OR interactions.