arXiv AI
Aug 25

Vision-Language Models for Occupational Physical Exposure Assessment: Estimating External Hand Forces in Manual Material Handling Tasks from RGB Video

The study presents a vision‑language model pipeline that estimates dynamic, triaxial, bilateral external hand forces during manual material handling tasks using only RGB video and known box mass. By combining text‑guided ROI localization, pretrained vision‑transformer features, and transformer‑based temporal regression, the model achieved root mean square errors of about 4.7–5.6 N for horizontal and mediolateral forces and 10.6–11.0 N for vertical forces across various camera setups. The approach demonstrated that including the handled object as a second ROI and using multi‑camera capture improved peak‑force estimation, showing the feasibility of sensor‑free force estimation for occupational exposure assessment.

By Mohammad Sadra Rajabi, Aanuoluwapo Ojelade, Sunwook Kim, Maury A. Nussbaum
arXiv Computer Vision
Sep 3

Video-Based Palm-Vein Authentication under Challenging Conditions

The paper introduces the Columbia University Palm‑vein (CUP) dataset, the first public video‑based palm‑vein dataset that captures palms under four surface conditions—clean, warm, wet, and dirty—along with physiological and demographic metadata. Twenty‑one recognizers are benchmarked on CUP, revealing that models performing well on clean palms lose most accuracy on dirty palms, with mean EER roughly quadrupling. The authors propose a lightweight design that fuses global cosine similarity with a saliency‑steered region‑level optimal transport, achieving state‑of‑the‑art performance across all surfaces while reducing parameters and computational cost, and they identify demographic gaps in warm‑condition performance.

By Xiaofeng Yan, Kechen Liu, Abhilash Venkatesh, Cathy Zhang, Xia Zhou, Salvatore Stolfo