arXiv Computer Vision

Occluded Gait Recognition with Mixture of Experts: An Action Detection Perspective

arXiv Computer Vision
Sep 11

MMGait: Benchmarking and Unifying Gait Recognition across Heterogeneous Modalities

MMGait is a large‑scale multi‑sensor benchmark that aligns visible, infrared, depth, LiDAR, and radar observations at the sequence level, enabling evaluation of single‑modal, cross‑modal, and multi‑modal gait recognition. The study shows that modality rankings shift with probe conditions, cross‑modal alignment remains challenging, and fusion can yield complementary gains. To address the scalability issue of training separate experts, the authors propose Omni‑Modal Gait Recognition and its implementation, OmniGait++, which unifies all recognition settings within a shared identity space using modality‑specific front ends, a shared encoder, and an anchor‑guided fusion module. whyItMatters":"MMGait provides a common testbed for heterogeneous gait sensing and demonstrates that unified recognition across varying modality availability is feasible, offering a scalable alternative to task‑specific experts."

By Saihui Hou, Chenye Wang, Qingyuan Cai, Aoqi Li, Yongzhen Huang
arXiv Computer Vision
Aug 28

Who Remains, What Changes: Identity Anchored Composed Gait Retrieval

The paper introduces Composed Gait Retrieval (CoGR), a task that retrieves a target gait sequence using a reference sequence and a natural language modification query. To support this, the authors create the first gait-language datasets—Language‑Augmented CCPG and Language‑Augmented CASIA‑B—via an automated annotation pipeline powered by large vision‑language models. They propose ComposeGait, an identity‑anchored composition framework with a Part‑aware Identity Adapter that injects identity tokens into a shared Q‑Former, achieving state‑of‑the‑art retrieval performance on both benchmarks.

By Jingchen Fei, Zengbin Wang, Yukun Liu, Muyi Sun, Shibiao Xu, Man Zhang
arXiv Computer Vision
1d ago

Vocabulary-Guided Gait Recognition

arXiv:2609.18413v1 Announce Type: new Abstract: What is a gait? Appearance-based gait networks consider a gait as the human shape and motion information from images. Model-based gait networks treat a...

By Panjian Huang, Saihui Hou, Chunshui Cao, Xu Liu, Yongzhen Huang
arXiv Computer Vision
Sep 10

3rd Place Solution to Human Motion Challenges in Real-World and Clinical Settings (MoCha) @ECCV2026: Language-Aligned Motion Representations for Domain-Generalizable UPDRS-Gait Severity Estimation

arXiv:2609.10187v1 Announce Type: new Abstract: In this work, we introduce language-aligned motion representations for domain-generalizable UPDRS-Gait severity estimation, aiming to learn semanticall...

By Soojie Kim, Muhammad Munsif, Minkyung Kim, Seungryul Baek
arXiv AI
Jun 2

From Human Videos to Robot Manipulation: A Survey on Scalable Vision-Language-Action Learning with Human-Centric Data

arXiv:2606. 00054v1 Announce Type: cross Abstract: Recent progress in generalizable embodied control has been driven by large-scale pretraining of Vision-Language-Action (VLA) models.

By Zhiyuan Feng, Qixiu Li, Huizhi Liang, Rushuai Yang, Yichao Shen, Zhiying Du, Zhaowei Zhang, Yu Deng, Li Zhao, Hao Zhao, Zongqing Lu, Oier Mees, Marc Pollefeys, Jiaolong Yang, Baining Guo