arXiv AI
Aug 25

Vision-Language Models for Occupational Physical Exposure Assessment: Estimating External Hand Forces in Manual Material Handling Tasks from RGB Video

The study presents a vision‑language model pipeline that estimates dynamic, triaxial, bilateral external hand forces during manual material handling tasks using only RGB video and known box mass. By combining text‑guided ROI localization, pretrained vision‑transformer features, and transformer‑based temporal regression, the model achieved root mean square errors of about 4.7–5.6 N for horizontal and mediolateral forces and 10.6–11.0 N for vertical forces across various camera setups. The approach demonstrated that including the handled object as a second ROI and using multi‑camera capture improved peak‑force estimation, showing the feasibility of sensor‑free force estimation for occupational exposure assessment.

By Mohammad Sadra Rajabi, Aanuoluwapo Ojelade, Sunwook Kim, Maury A. Nussbaum
arXiv Machine Learning
Jun 29

Cross-view Multimodal Vision-Based Assessment Framework for Traditional Chinese Medicine Rehabilitation Training

arXiv:2606. 28104v1 Announce Type: cross Abstract: Vision-based assessment can provide convenient and cost-effective evaluation in Traditional Chinese Medicine (TCM) rehabilitation training, where action quality assessment (AQA) from computer vision offers a promising solution.

By Francis Xiatian Zhang, Hao Yao, Shengxuan Chen, Hong Zhu, Hongxiao Jia, Sisi Zheng, Hubert P. H. Shum
Hugging Face Trending Papers
Jun 3

VISTA: Vision-Grounded and Physics-Validated Adaptation of UMI data for VLA Training

Universal Manipulation Interface (UMI) enables scalable real-world robot data collection without hardware-specific teleoperation, yet leveraging UMI data to train large-scale Vision-Language-Action (VLA) models remains fundamentally challenging. We identify two critical mismatches: wrist-mounted fisheye views, with severe radial distortion and local gripper-centric perspectives, are out-of-distribution for pretrained VLMs; and human-collected trajectories frequently violate kinematic limits, incur collisions, or exceed controller bandwidth, teaching VLA policies physically infeasible actions.