The study investigates how well a single consumer earbud IMU can estimate 3D body pose and whether adding foot IMUs improves accuracy. Using a multimodal capture pipeline with RGB‑D video, an AirPods head IMU, and Striv insole IMUs, the authors benchmark pose estimation across various motions and train recurrent models (IMUPoser and MobilePoser). Results show that a head IMU alone achieves 79.0 mm rigid‑MPJPE and 0.809 macro‑F1 for foot contact, while adding foot IMUs does not significantly improve pose and can even degrade performance due to insole orientation errors.
arXiv:2609.22619v1 Announce Type: new
Abstract: Tracking recovery of walking function requires detecting meaningful gait change across rehabilitation sessions, yet objective 3D measurement remains co...
By Nethmi Jayasinghe, Mihir Parashar, Amit Ranjan Trivedi
MuyBridge is an on-device system that estimates an athlete’s segmental center of mass (CoM) trajectory from a single phone camera video stream. It combines a compact 2D pose network with a distilled monocular depth network, fusing their outputs through anatomical and physical priors to produce metric CoM estimates without requiring 3D or task‑specific supervision. On the AthletePose3D dataset, MuyBridge achieves 33–41 mm vertical CoM error and 2.3–6.6 % absolute‑relative range error, delivering CoM estimates at 63 FPS on an iPhone 15 with asynchronous depth updates.
By Aidan Bradshaw, Marco Giordano, David Rode, Andreas Habersack, Elif Basokur, Annika Kruse, Markus Tilp, Michele Magno, Peter Wolf, Luca Benini, Christoph Leitner
The paper evaluates six onboarding strategies for federated wearable models on five datasets using a leakage‑controlled protocol that fixes source checkpoints and separates calibration from evaluation. Results show that while average accuracy is high, person‑level performance can drop significantly, with some methods causing negative transfer for certain users. The study highlights that mean accuracy alone is insufficient and provides an auditable benchmark and failure map for future development.
By Rahil Aftab, Vineet Kumar Rakesh, Soumya Mazumdar, Tapas Samanta
The paper presents a method for improving markerless 3D pose estimation in infants by cross‑model distillation. Using unannotated infant video, a frozen Sapiens 2 pose model provides dense pseudo‑labels that guide fine‑tuning of the SAM 3D Body model. On a held‑out dataset of eleven infants, the fine‑tuned model shows significant gains in 2D keypoint accuracy and 3D joint error compared to the original SAM 3D Body model.
By R. James Cotton, Divya Joshi, Colleen Peyton
arXiv:2609.09670v1 Announce Type: new
Abstract: Monocular pose estimation enables low-cost gait analysis but is sensitive to missing keypoints caused by occlusion, detection errors, or efficiency-dri...
By Shubham Jariwala
arXiv:2606. 31127v1 Announce Type: cross Abstract: To enable personalized, real-time coaching using Augmented Reality glasses or fixed camera setups in domains such as sports, cooking, or music, a system must understand not just what a person does, but how well they execute an activity.
By Bj\"orn Braun, Christian Holz
arXiv:2608. 19480v1 Announce Type: new Abstract: Human pose estimation has advanced significantly due to the development of deep learning models, increased data availability, and improved computing resources.
By Luis F. Gomez, Julian Fierrez, Roberto Daza, Ruben Tolosana, Aythami Morales, Gonzalo Garrido, Javier Rueda, Enrique Navarro
arXiv:2609.37297v1 Announce Type: new
Abstract: Cross-skeleton motion generation trains generative models to carry action structure and motion intention from one body to another. Yet a target motion...
By Zhiyuan Li, Wenyan Yang, Pekka Marttinen, Joni Pajarinen
arXiv:2608.29928v1 Announce Type: new
Abstract: State-of-the-art monocular body recovery methods predict mesh vertices and angles on the corresponding kinematic tree, but their outputs lack biomechan...
By R. James Cotton, J. D. Peiffer, Lucinda Williamson, John Leske, Georgios Pavlakos
The paper introduces the Columbia University Palm‑vein (CUP) dataset, the first public video‑based palm‑vein dataset that captures palms under four surface conditions—clean, warm, wet, and dirty—along with physiological and demographic metadata. Twenty‑one recognizers are benchmarked on CUP, revealing that models performing well on clean palms lose most accuracy on dirty palms, with mean EER roughly quadrupling. The authors propose a lightweight design that fuses global cosine similarity with a saliency‑steered region‑level optimal transport, achieving state‑of‑the‑art performance across all surfaces while reducing parameters and computational cost, and they identify demographic gaps in warm‑condition performance.
By Xiaofeng Yan, Kechen Liu, Abhilash Venkatesh, Cathy Zhang, Xia Zhou, Salvatore Stolfo
arXiv:2202. 14019v3 Announce Type: replace-cross Abstract: Maintaining proper form while exercising is important for preventing injuries and maximizing muscle mass gains.
By Paritosh Parmar, Amol Gharat, Helge Rhodin