arXiv AI

HandAnthro: Automated Hand Anthropometry from a Single Image

arXiv AI
Aug 25

Vision-Language Models for Occupational Physical Exposure Assessment: Estimating External Hand Forces in Manual Material Handling Tasks from RGB Video

The study presents a vision‑language model pipeline that estimates dynamic, triaxial, bilateral external hand forces during manual material handling tasks using only RGB video and known box mass. By combining text‑guided ROI localization, pretrained vision‑transformer features, and transformer‑based temporal regression, the model achieved root mean square errors of about 4.7–5.6 N for horizontal and mediolateral forces and 10.6–11.0 N for vertical forces across various camera setups. The approach demonstrated that including the handled object as a second ROI and using multi‑camera capture improved peak‑force estimation, showing the feasibility of sensor‑free force estimation for occupational exposure assessment.

By Mohammad Sadra Rajabi, Aanuoluwapo Ojelade, Sunwook Kim, Maury A. Nussbaum
arXiv Computer Vision
Sep 3

Video-Based Palm-Vein Authentication under Challenging Conditions

The paper introduces the Columbia University Palm‑vein (CUP) dataset, the first public video‑based palm‑vein dataset that captures palms under four surface conditions—clean, warm, wet, and dirty—along with physiological and demographic metadata. Twenty‑one recognizers are benchmarked on CUP, revealing that models performing well on clean palms lose most accuracy on dirty palms, with mean EER roughly quadrupling. The authors propose a lightweight design that fuses global cosine similarity with a saliency‑steered region‑level optimal transport, achieving state‑of‑the‑art performance across all surfaces while reducing parameters and computational cost, and they identify demographic gaps in warm‑condition performance.

By Xiaofeng Yan, Kechen Liu, Abhilash Venkatesh, Cathy Zhang, Xia Zhou, Salvatore Stolfo
arXiv AI
Aug 12

A HamNoSys-Guided Dataset and Baselines for Fine-Grained Isolated Handshape Recognition in Sign Language

arXiv:2608. 10588v1 Announce Type: cross Abstract: Purpose: Fine-grained handshape recognition supports computational sign-language transcription, recognition, and translation, but broad, phonetically defined visual inventories with signer-aware evaluation remain limited.

By Ushnish Sarkar, Suvajit Patra, Bhaswar Chattopadhyay, Pranab Singha Roy, Tapas Samanta
arXiv Computer Vision
Sep 14

KAD-Net: Kinematics-Aware Decoupled Learning for Robust 3D Hand Pose Estimation from a Single Depth Image

KAD-Net introduces a Kinematics-Aware Decoupled Learning Network for 3D hand pose estimation from a single depth image. It employs a Finger Topology Constraint module that uses local kinematic representations of three consecutive finger joints to better model distal joint relationships and handle occlusion. The architecture also decouples 2D joint localization from depth estimation in a hierarchical multitask framework, reducing feature interference and improving accuracy on benchmark datasets such as ICVL, NYU, and MSRA.

By Jun Lu, Zhenming Chen, Lin Chen, Kanlun Tan, Xiaoling Li, Qiao Liu
arXiv Machine Learning
Jun 29

Cross-view Multimodal Vision-Based Assessment Framework for Traditional Chinese Medicine Rehabilitation Training

arXiv:2606. 28104v1 Announce Type: cross Abstract: Vision-based assessment can provide convenient and cost-effective evaluation in Traditional Chinese Medicine (TCM) rehabilitation training, where action quality assessment (AQA) from computer vision offers a promising solution.

By Francis Xiatian Zhang, Hao Yao, Shengxuan Chen, Hong Zhu, Hongxiao Jia, Sisi Zheng, Hubert P. H. Shum