WildHSR introduces a lightweight adaptation of 3D foundation models to jointly recover metric cameras, scene geometry, and persistent person identities from monocular video. By generating pseudo‑scale labels from curated web footage and fine‑tuning a Scale Readout, the method predicts metric scale directly from foundation‑model tokens. It also exploits intermediate query‑key features to associate per‑frame bodies, enabling feed‑forward reconstruction that outperforms state‑of‑the‑art optimization‑based methods on several benchmarks while running at 10.1 fps.
By Jerrin Bright, John Zelek
arXiv:2607. 17342v1 Announce Type: cross Abstract: Understanding physical human-robot and human-human interactions is a challenging yet emerging topic in 3D vision.
By Yuhang Wen, Mengyuan Liu, Zixuan Tang, Junsong Yuan, Sirui Li, Beichen Ding
arXiv:2608.29928v1 Announce Type: new
Abstract: State-of-the-art monocular body recovery methods predict mesh vertices and angles on the corresponding kinematic tree, but their outputs lack biomechan...
By R. James Cotton, J. D. Peiffer, Lucinda Williamson, John Leske, Georgios Pavlakos
Driven by the availability of large-scale datasets, Human Pose Estimation (HPE) plays a critical role in numerous downstream tasks. However, mainstream benchmarks exhibit severe representation bias, predominantly featuring able-bodied individuals.
arXiv:2609.15669v1 Announce Type: cross
Abstract: Multimodal image registration is a key component of many clinical workflows, yet it remains challenging because corresponding anatomical structures o...
By Matteo Barbieri, Giammarco La Barbera, Juan Pablo De La Plata, Sabine Sarnacki, Isabelle Bloch, Pietro Gori
arXiv:2609.18406v1 Announce Type: new
Abstract: Recovering 3D human body motion from video is important for applications such as rehabilitation assessment and sports performance evaluation. For prost...
By Yilin Wen, Kechuan Dong, Fumiya Suginaka, Ken Endo, Yusuke Sugano
TokenMatch is a transformer-based model that estimates 3D shape correspondences by adaptively tokenising meshes into curvature-guided patches. Trained only on the BeCoS partial-to-partial dataset, it generalises to full-shape matching without retraining, using self‑ and cross‑attention to learn patch‑ and point‑level relations. Evaluated on CP2P, PSMAL, BeCoS, FAUST, SCAPE, and SHREC'19, TokenMatch consistently outperforms existing methods in mean geodesic error and intersection‑over‑union while achieving sub‑second inference speeds.
By Adeela Islam, Zorah L\"ahner, Vittorio Murino, Vladislav Golyanik
Human mesh recovery (HMR) aims to recover 3D human meshes from images. Most existing HMR benchmarks and methods focus on either multi-person reconstruction from a single view or single-person reconstruction from multiple views, where the number of subjects and the scene scale are relatively limited.
arXiv:2601.13913v3 Announce Type: replace
Abstract: We consider monocular 3D human pose estimation (HPE), where the goal is to predict 3D human skeletal joints from a single 2D image, typically via 2...
By Pavlo Melnyk, Cuong Le, Urs Waldmann, Per-Erik Forss\'en, Bastian Wandt
arXiv:2607. 13905v1 Announce Type: cross Abstract: The International StepUP Competition Series was launched to advance research in pressure-based footstep biometrics through a standardized and challenging evaluation framework.
By Robyn Larracy, Anant Gupta, Gourav Gupta, Ethan Eddy, Maxime Devanne, Cyril Meyer, Jin-Chern Chiou, Yueh-Shan Lee, Zong-Han Lu, Aaron Tabor, Erik Scheme
arXiv:2607. 08725v1 Announce Type: cross Abstract: Recent progress in 3D human pose estimation has made markerless recovery of skeletal motion increasingly accurate and scalable.
By Ayda Eghbalian, Kevin Desai
Non-rigid 3D shape matching is a fundamental task in computer vision and graphics. In this paper, we propose a hybrid self-supervised method based on a coarse-to-fine strategy, which ensures consistency between the coarse mapping and the refined correspondence produced by our refinement module.