arXiv:2606. 01896v1 Announce Type: cross Abstract: Generated (or synthetic) image data is increasingly used to augment or replace real training datasets when target imagery is scarce, expensive, or biased.
By Atmika Bhardwaj, Silvia Vock, Nico Steckhan
arXiv:2606. 25619v2 Announce Type: replace Abstract: In this paper, we present ScaleHP, a unified framework that explicitly represents per-instance metric scale to resolve the coupled errors in calibrated camera-space hand pose estimation.
By Ruitao Jing, Xingyu Chen, Hongyang Li, Qing Jiang, Yukai Shi, Lei Zhang
External hand forces are important inputs to biomechanical analyses of occupational physical exposure and injury risk, yet continuous force measurements during manual material handling (MMH) typically...
The study presents a vision‑language model pipeline that estimates dynamic, triaxial, bilateral external hand forces during manual material handling tasks using only RGB video and known box mass. By combining text‑guided ROI localization, pretrained vision‑transformer features, and transformer‑based temporal regression, the model achieved root mean square errors of about 4.7–5.6 N for horizontal and mediolateral forces and 10.6–11.0 N for vertical forces across various camera setups. The approach demonstrated that including the handled object as a second ROI and using multi‑camera capture improved peak‑force estimation, showing the feasibility of sensor‑free force estimation for occupational exposure assessment.
By Mohammad Sadra Rajabi, Aanuoluwapo Ojelade, Sunwook Kim, Maury A. Nussbaum
The paper introduces the Columbia University Palm‑vein (CUP) dataset, the first public video‑based palm‑vein dataset that captures palms under four surface conditions—clean, warm, wet, and dirty—along with physiological and demographic metadata. Twenty‑one recognizers are benchmarked on CUP, revealing that models performing well on clean palms lose most accuracy on dirty palms, with mean EER roughly quadrupling. The authors propose a lightweight design that fuses global cosine similarity with a saliency‑steered region‑level optimal transport, achieving state‑of‑the‑art performance across all surfaces while reducing parameters and computational cost, and they identify demographic gaps in warm‑condition performance.
By Xiaofeng Yan, Kechen Liu, Abhilash Venkatesh, Cathy Zhang, Xia Zhou, Salvatore Stolfo
arXiv:2609.24424v1 Announce Type: new
Abstract: Monocular RGB-based hand pose estimation has emerged as a critical research frontier in computer vision. The local hand pose estimation methods predict...
By Kaiwen Ren, Yiran Jiang, Yongjing Ye, Shihong Xia
arXiv:2608. 10588v1 Announce Type: cross Abstract: Purpose: Fine-grained handshape recognition supports computational sign-language transcription, recognition, and translation, but broad, phonetically defined visual inventories with signer-aware evaluation remain limited.
By Ushnish Sarkar, Suvajit Patra, Bhaswar Chattopadhyay, Pranab Singha Roy, Tapas Samanta
KAD-Net introduces a Kinematics-Aware Decoupled Learning Network for 3D hand pose estimation from a single depth image. It employs a Finger Topology Constraint module that uses local kinematic representations of three consecutive finger joints to better model distal joint relationships and handle occlusion. The architecture also decouples 2D joint localization from depth estimation in a hierarchical multitask framework, reducing feature interference and improving accuracy on benchmark datasets such as ICVL, NYU, and MSRA.
By Jun Lu, Zhenming Chen, Lin Chen, Kanlun Tan, Xiaoling Li, Qiao Liu
arXiv:2606. 28104v1 Announce Type: cross Abstract: Vision-based assessment can provide convenient and cost-effective evaluation in Traditional Chinese Medicine (TCM) rehabilitation training, where action quality assessment (AQA) from computer vision offers a promising solution.
By Francis Xiatian Zhang, Hao Yao, Shengxuan Chen, Hong Zhu, Hongxiao Jia, Sisi Zheng, Hubert P. H. Shum
arXiv:2609.13269v1 Announce Type: cross
Abstract: Gesture recognition on video is normally posed as classification: label each frame, then act on the label. That is adequate for control, where a comm...
By Amey Thakur
Hand Pose Estimation (HPE) is a fundamental technology for various applications such as AR/VR and robotics. In these applications, the visibility of each hand joint in the image is crucial for assessing the reliability of estimation results under occlusion.
Quantitative joint angles are rarely available in routine care because the tools are slow, costly, or confined to a laboratory. We show that clinical joint angles can be read directly from the per-segment rotation matrices a parametric body model already produces, with no inverse-kinematics or musculoskeletal-model fitting step.