arXiv:2609.18490v1 Announce Type: new
Abstract: "What I cannot create, I do not understand."Human wisdom reveals that creation is one of the highest forms of learning. For example, Diffusion Models h...
By Panjian Huang, Saihui Hou, Junzhou Huang, Yongzhen Huang
arXiv:2609.18432v1 Announce Type: new
Abstract: Extensive occlusions in real-world scenarios pose challenges to gait recognition due to missing and noisy information, as well as body misalignment in...
By Panjian Huang, Yunjie Peng, Saihui Hou, Chunshui Cao, Xu Liu, Zhiqiang He, Yongzhen Huang
arXiv:2512. 00691v2 Announce Type: replace Abstract: Gait patterns play a critical role in human identification and healthcare analytics, yet current progress remains constrained by small, narrowly designed models that fail to scale or generalize.
By Dingqiang Ye, Chao Fan, Kartik Narayan, Bingzhe Wu, Chengwen Luo, Jianqiang Li, Vishal M. Patel
The paper introduces Composed Gait Retrieval (CoGR), a task that retrieves a target gait sequence using a reference sequence and a natural language modification query. To support this, the authors create the first gait-language datasets—Language‑Augmented CCPG and Language‑Augmented CASIA‑B—via an automated annotation pipeline powered by large vision‑language models. They propose ComposeGait, an identity‑anchored composition framework with a Part‑aware Identity Adapter that injects identity tokens into a shared Q‑Former, achieving state‑of‑the‑art retrieval performance on both benchmarks.
By Jingchen Fei, Zengbin Wang, Yukun Liu, Muyi Sun, Shibiao Xu, Man Zhang
arXiv:2609.10187v1 Announce Type: new
Abstract: In this work, we introduce language-aligned motion representations for domain-generalizable UPDRS-Gait severity estimation, aiming to learn semanticall...
By Soojie Kim, Muhammad Munsif, Minkyung Kim, Seungryul Baek
MMGait is a large‑scale multi‑sensor benchmark that aligns visible, infrared, depth, LiDAR, and radar observations at the sequence level, enabling evaluation of single‑modal, cross‑modal, and multi‑modal gait recognition. The study shows that modality rankings shift with probe conditions, cross‑modal alignment remains challenging, and fusion can yield complementary gains. To address the scalability issue of training separate experts, the authors propose Omni‑Modal Gait Recognition and its implementation, OmniGait++, which unifies all recognition settings within a shared identity space using modality‑specific front ends, a shared encoder, and an anchor‑guided fusion module.
whyItMatters":"MMGait provides a common testbed for heterogeneous gait sensing and demonstrates that unified recognition across varying modality availability is feasible, offering a scalable alternative to task‑specific experts."
By Saihui Hou, Chenye Wang, Qingyuan Cai, Aoqi Li, Yongzhen Huang
arXiv:2508.13073v3 Announce Type: replace-cross
Abstract: Robotic manipulation, a key frontier in robotics and embodied AI, requires precise motor control and multimodal understanding, yet traditiona...
By Rui Shao, Wei Li, Lingsen Zhang, Renshan Zhang, Zhiyang Liu, Ran Chen, Liqiang Nie
VKnowU is a benchmark that tests multimodal large language models (MLLMs) on their grasp of visual knowledge—intuitive, human-like understanding of physical and social principles in videos. The benchmark contains 1,680 questions across 1,249 videos, covering eight core types of visual knowledge, and shows that current state‑of‑the‑art MLLMs still lag behind human performance, especially on world‑centric tasks. To address this gap, the authors release VKnowQA and VideoKnow+, a baseline model that incorporates visual knowledge via a See‑Think‑Answer framework and reinforcement learning, improving performance on VKnowU and related datasets.
By Tianxiang Jiang, Sheng Xia, Yicheng Xu, Linquan Wu, Xiangyu Zeng, Limin Wang, Yu Qiao, Yi Wang
arXiv:2605. 16713v2 Announce Type: replace-cross Abstract: Modern Vision-Language Models (VLMs) achieve strong semantic recognition, yet remain brittle on elementary spatial relations such as left of, on, behind, and between.
By Renjie Gu, Kaichen Zhou, Yan Luo, Mengyu Wang
arXiv:2609.08038v2 Announce Type: cross
Abstract: Smart healthcare monitoring systems require precise action recognition to ensure well-being and timely intervention in critical situations such as fa...
By Diwas Lamsal, Pramod Wickramatilake, Jednipat Moonrinta, Mongkol Ekpanyapong, Matthew N. Dailey
arXiv:2608. 10932v1 Announce Type: cross Abstract: Understanding camera motion is fundamental to video perception, with applications in spatial intelligence and controllable video generation.
By Dazhao Du, Shiyan Du, Jian Liu, Yongjian Yu, Bohai Gu, Tao Han, Hualuo Liu, Eric Liu, Yujia Zhang, Xi Chen, Song Guo
arXiv:2606. 08530v1 Announce Type: cross Abstract: Vision-Language-Action (VLA) models achieve strong benchmark performance but still struggle in real-world deployment with unseen objects, background shifts, and different robot embodiments.
By Yuan Zhang, Shiqi Zhang, Yedong Shen, Shuai Dong, Jiajun Deng, Xin Zhang, Yuxuan Gao, Jiajia Wu, Xin Nie, Zhiyuan Cheng, Jianmin Ji, Yanyong Zhang, Xingyi Zhang, Jia Pan