arXiv Computer Vision By R. James Cotton, Divya Joshi, Colleen Peyton

Cross-Model Distillation of a Human-Pose Foundation Model from Unannotated Infant Video for Markerless 3D Pose Estimation

Read the original on arXiv Computer Vision →

The paper presents a method for improving markerless 3D pose estimation in infants by cross‑model distillation. Using unannotated infant video, a frozen Sapiens 2 pose model provides dense pseudo‑labels that guide fine‑tuning of the SAM 3D Body model. On a held‑out dataset of eleven infants, the fine‑tuned model shows significant gains in 2D keypoint accuracy and 3D joint error compared to the original SAM 3D Body model.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Computer Vision.

arXiv Machine Learning
Jun 29

Cross-view Multimodal Vision-Based Assessment Framework for Traditional Chinese Medicine Rehabilitation Training

arXiv:2606. 28104v1 Announce Type: cross Abstract: Vision-based assessment can provide convenient and cost-effective evaluation in Traditional Chinese Medicine (TCM) rehabilitation training, where action quality assessment (AQA) from computer vision offers a promising solution.

By Francis Xiatian Zhang, Hao Yao, Shengxuan Chen, Hong Zhu, Hongxiao Jia, Sisi Zheng, Hubert P. H. Shum
Hugging Face Trending Papers
Aug 4

Learning Biomechanically Plausible Human Motion from Sparse Radar Point Clouds

Radar-based human pose estimation has focused on improving learning algorithms while representing the body as unconstrained keypoint coordinates. We address the underexplored dimension of anatomical fidelity by integrating a full-body skeletal model into a differentiable, end-to-end trainable radar-based pose estimation framework, in which the pose network is supervised through forward kinematics while subject-specific geometry is fitted beforehand.