Hugging Face Trending Papers

One Sensor, Whole Body - 3D Body Pose from a Single Consumer Earbud IMU

The study investigates how well a single consumer earbud IMU can estimate 3D body pose and whether adding foot IMUs improves accuracy. Using a multimodal capture pipeline with RGB‑D video, an AirPods head IMU, and Striv insole IMUs, the authors benchmark pose estimation across various motions and train recurrent models (IMUPoser and MobilePoser). Results show that a head IMU alone achieves 79.0 mm rigid‑MPJPE and 0.809 macro‑F1 for foot contact, while adding foot IMUs does not significantly improve pose and can even degrade performance due to insole orientation errors.

arXiv Computer Vision
Sep 3

MuyBridge: Mobile Human Center-of-Mass Estimation from Monocular Video via Sparse Fusion

MuyBridge is an on-device system that estimates an athlete’s segmental center of mass (CoM) trajectory from a single phone camera video stream. It combines a compact 2D pose network with a distilled monocular depth network, fusing their outputs through anatomical and physical priors to produce metric CoM estimates without requiring 3D or task‑specific supervision. On the AthletePose3D dataset, MuyBridge achieves 33–41 mm vertical CoM error and 2.3–6.6 % absolute‑relative range error, delivering CoM estimates at 63 FPS on an iPhone 15 with asynchronous depth updates.

By Aidan Bradshaw, Marco Giordano, David Rode, Andreas Habersack, Elif Basokur, Annika Kruse, Markus Tilp, Michele Magno, Peter Wolf, Luca Benini, Christoph Leitner
arXiv Computer Vision
Sep 17

Video-Based Markerless Motion Capture for Clinical and Rehabilitation Biomechanics: A PRISMA-ScR Scoping Review of Validated Architectures, Clinical Readiness, and Emerging Methods

This scoping review examined 117 studies on video-based markerless motion capture, most published from 2024 onward and focused on healthy adults walking in laboratories. The studies identified five main pipeline architectures, but most reported only raw joint angles without biomechanical refinement, achieving sagittal lower‑limb agreement of about 5–6°, which falls short of clinical acceptability. Validation of out‑of‑plane kinematics, kinetics, and performance in older or pathological populations was rare, and emerging computer‑vision techniques such as foundation‑model mesh recovery and differentiable inverse kinematics were largely absent from validated work.

By Florian Delaplace (LAMHESS, CHU), Elodie Piche (LAMHESS), Fr\'ed\'eric Chorin (IUF, LAMHESS), Raphael Zory (IUF, LAMHESS)
arXiv Computer Vision
Sep 3

Cross-Model Distillation of a Human-Pose Foundation Model from Unannotated Infant Video for Markerless 3D Pose Estimation

The paper presents a method for improving markerless 3D pose estimation in infants by cross‑model distillation. Using unannotated infant video, a frozen Sapiens 2 pose model provides dense pseudo‑labels that guide fine‑tuning of the SAM 3D Body model. On a held‑out dataset of eleven infants, the fine‑tuned model shows significant gains in 2D keypoint accuracy and 3D joint error compared to the original SAM 3D Body model.

By R. James Cotton, Divya Joshi, Colleen Peyton
arXiv Computer Vision
Sep 25

Training-Free Hold-Usage Detection in Sport Climbing with Foundation Pose Models

The paper presents a training‑free method for detecting which holds a climber uses in sport climbing videos by leveraging a frozen foundation pose model (Sapiens) that provides fingertip and toe keypoints. Using a simple proximity test, mutual exclusion, and a temporal‑persistence rule, the approach achieves high F_1 scores (up to 90.2%) on the Way Up dataset without any climbing‑specific training, outperforming repurposed pose pipelines. The resulting automatic predictions enable accurate coaching statistics, such as climb time and pace, with Pearson correlations of 1.00 and 0.94 respectively.

By Abu Bakar, Abdullah Aftab, Amir Hamza
Hugging Face Trending Papers
Aug 4

Learning Biomechanically Plausible Human Motion from Sparse Radar Point Clouds

Radar-based human pose estimation has focused on improving learning algorithms while representing the body as unconstrained keypoint coordinates. We address the underexplored dimension of anatomical fidelity by integrating a full-body skeletal model into a differentiable, end-to-end trainable radar-based pose estimation framework, in which the pose network is supervised through forward kinematics while subject-specific geometry is fitted beforehand.