Hugging Face Trending Papers

Partial Skeleton Visibility for Action Recognition: A Constrained Field-of-View Approach

Skeleton-based action recognition has achieved remarkable success by exploiting joint coordinates and their topological connections, yet prevailing methods overwhelmingly assume complete and clean skeleton inputs. In real-world deployments, such as egocentric vision, crowded surveillance, wearable devices, or edge robotics, limited field-of-view (FoV) frequently causes substantial joint visibility dropout, leading to severe performance degradation that existing models are largely unprepared to handle.

arXiv AI
Sep 18

FAMOS: Feed-Forward 3D Articulation Modeling from Sparse Observations

FAMOS is a feed‑forward model that predicts movable‑part segmentation and joint parameters from a sparse, unordered set of partial point clouds. It jointly reasons over multiple observations using a Multi‑state Articulation Transformer that alternates state‑wise and global attention, and introduces an observed articulation span objective to supervise motion ranges across inputs. A procedural data generator supplies self‑annotated assets for training, and experiments on PartNet‑Mobility, ACD, and ArtiCraft‑10K show consistent improvements over existing feed‑forward and optimization‑based baselines.

By Kevin Qu, Tao Sun, Massimiliano Viola, Liyuan Zhu, Zhizhuo Zhou, Sayan Deb Sarkar, Konrad Schindler, Iro Armeni
arXiv Computer Vision
Sep 11

Artic-O: End-to-End Articulated Object Reconstruction via Latent Geometry Learning

Artic-O is an end‑to‑end, feed‑forward framework that reconstructs articulated objects from sparse images by learning latent geometry. It maps multi‑state observations into a pretrained latent geometry space, uses a frozen flow‑matching decoder for complete‑shape priors, and fuses visual tokens with geometry latents in an image‑grounded part‑reasoning module to segment active parts and predict articulation. Trained with a geometry‑to‑articulation curriculum and a decoupled two‑pass strategy, Artic‑O achieves high reconstruction quality and articulation accuracy while drastically reducing inference time from 9 minutes to about 0.3 seconds per object.

By Xuyang Wang, Zhenyu Li, Jian Ding, Habib Slim, Peter Wonka, Hongdong Li, Mohamed Elhoseiny
arXiv Computer Vision
Sep 15

MoCapAnything V2: End-to-End Motion Capture for Arbitrary Skeletons

arXiv:2604.28130v4 Announce Type: replace Abstract: Recent methods for arbitrary-skeleton motion capture from monocular video follow a factorized pipeline, where a Video-to-Pose network predicts join...

By Kehong Gong, Zhengyu Wen, Dao Thien Phong, Mingxi Xu, Weixia He, Qi Wang, Ning Zhang, Zhengyu Li, Guanli Hou, Dongze Lian, Xiaoyu He, Mingyuan Zhang, Hanwang Zhang
arXiv Computer Vision
Aug 27

Pose-Anchored Optical Flow for Low-Latency Human Action Anticipation in Human-Robot Teaming

Pose-Anchored Optical Flow for Low-Latency Human Action Anticipation in Human-Robot Teaming proposes PoseOFF, a representation that captures local motion around human joints by conditioning optical flow extraction on pose. This structured motion representation aligns with human kinematics and improves early action recognition accuracy across multiple datasets and backbones. PoseOFF achieves comparable or better performance while observing less of the action sequence, making it suitable for real‑time, resource‑constrained robotic systems.

By Lewis de Zoete Grundy, Chris McCarthy, Christopher Fluke
arXiv Computer Vision
Sep 18

DirtyMoCap: Robust Motion Capture from Unconstrained Markers

DirtyMoCap is a marker‑layout‑free framework that converts unordered, noisy optical motion capture markers into a fixed set of proxy anchors representing skeletal joints and body surface points. Using a recurrent sliding‑window architecture to track these anchors and a custom differentiable Gauss‑Newton solver to fit the SMPL‑H model, the method learns adaptive observation confidence, smoothness, and prior weights end‑to‑end. Experiments show that DirtyMoCap generalizes across arbitrary marker configurations, outperforms configuration‑specific baselines in joint and vertex accuracy, and achieves up to a 100× speedup over standard PyTorch implementations, enabling the creation of a temporally coherent Kung Fu motion dataset.

By Long Wang, Shuting Zhao, Shen Yan, Siyuan Yu, Xiaoben Li, Zeyu Cai, Yumeng Hou, Yuliang Xiu