Skeleton-based action recognition has achieved remarkable success by exploiting joint coordinates and their topological connections, yet prevailing methods overwhelmingly assume complete and clean skeleton inputs. In real-world deployments, such as egocentric vision, crowded surveillance, wearable devices, or edge robotics, limited field-of-view (FoV) frequently causes substantial joint visibility dropout, leading to severe performance degradation that existing models are largely unprepared to handle.
arXiv:2605.14854v3 Announce Type: replace-cross
Abstract: Human Mesh Recovery (HMR) is fundamentally ambiguous: under occlusion or weak depth cues, multiple 3D bodies can explain the same image evide...
By Patrick Kwon, Chen Chen
arXiv:2607. 17342v1 Announce Type: cross Abstract: Understanding physical human-robot and human-human interactions is a challenging yet emerging topic in 3D vision.
By Yuhang Wen, Mengyuan Liu, Zixuan Tang, Junsong Yuan, Sirui Li, Beichen Ding
FAMOS is a feed‑forward model that predicts movable‑part segmentation and joint parameters from a sparse, unordered set of partial point clouds. It jointly reasons over multiple observations using a Multi‑state Articulation Transformer that alternates state‑wise and global attention, and introduces an observed articulation span objective to supervise motion ranges across inputs. A procedural data generator supplies self‑annotated assets for training, and experiments on PartNet‑Mobility, ACD, and ArtiCraft‑10K show consistent improvements over existing feed‑forward and optimization‑based baselines.
By Kevin Qu, Tao Sun, Massimiliano Viola, Liyuan Zhu, Zhizhuo Zhou, Sayan Deb Sarkar, Konrad Schindler, Iro Armeni
arXiv:2604.28130v4 Announce Type: replace
Abstract: Recent methods for arbitrary-skeleton motion capture from monocular video follow a factorized pipeline, where a Video-to-Pose network predicts join...
By Kehong Gong, Zhengyu Wen, Dao Thien Phong, Mingxi Xu, Weixia He, Qi Wang, Ning Zhang, Zhengyu Li, Guanli Hou, Dongze Lian, Xiaoyu He, Mingyuan Zhang, Hanwang Zhang
Artic-O is an end‑to‑end, feed‑forward framework that reconstructs articulated objects from sparse images by learning latent geometry. It maps multi‑state observations into a pretrained latent geometry space, uses a frozen flow‑matching decoder for complete‑shape priors, and fuses visual tokens with geometry latents in an image‑grounded part‑reasoning module to segment active parts and predict articulation. Trained with a geometry‑to‑articulation curriculum and a decoupled two‑pass strategy, Artic‑O achieves high reconstruction quality and articulation accuracy while drastically reducing inference time from 9 minutes to about 0.3 seconds per object.
By Xuyang Wang, Zhenyu Li, Jian Ding, Habib Slim, Peter Wonka, Hongdong Li, Mohamed Elhoseiny
arXiv:2609.24482v1 Announce Type: new
Abstract: Monocular 3D human pose estimation (HPE) remains challenging due to depth ambiguity, occlu- sions, and the need for temporal consistency. While multi-v...
By Mena Kamel, Natalie Won, Amrut Sarangi, Sven Jager, Albert Pla Planas
Fine-grained understanding of operating room (OR) activity could enable workflow-aware assistance, yet remains difficult due to clutter, occlusions, and limited sensing. The prevailing approach to model this environment is scene graphs as an interpretable representation of OR interactions.
arXiv:2608. 12187v1 Announce Type: cross Abstract: Transformer-based methods have achieved strong performance in monocular 3D human pose estimation, but most existing approaches organise spatial and temporal reasoning as separate stages, which may weaken unified spatial-temporal interdependencies inherent in human motion and compress frame-level structural information before temporal modelling.
By Ruochen Li, Shuang Chen, Wenke E, Farshad Arvin, Amir Atapour-Abarghouei
arXiv:2604. 09063v3 Announce Type: replace-cross Abstract: Human action recognition is pivotal in computer vision, with applications ranging from surveillance to human-robot interaction.
By Yuxi Zhou, Zhengbo Zhang, Jingyu Pan, Zhiyu Lin, Zhigang Tu
The paper introduces GaitMoE, an action‑detection based mixture‑of‑experts framework for occluded gait recognition, leveraging temporal and action experts to infer missing body parts from adjacent frames and gait cycles. It also presents a new Occluded Gait database (OccGait) with diverse occlusion scenarios and annotations, and demonstrates superior performance on OccGait, OccCASIA‑B, Gait3D, and GREW datasets.
By Panjian Huang, Yunjie Peng, Saihui Hou, Chunshui Cao, Xu Liu, Zhiqiang He, Yongzhen Huang
arXiv:2608. 19693v1 Announce Type: cross Abstract: Sparse keypoint extraction and matching underpin core tasks in geometric computer vision, including structure-from-motion, visual SLAM, augmented reality, and medical image registration.
By Johannes K\"unzel, Peter Eisert, Anna Hilsmann