arXiv Computer Vision By Wenjin Fu, Li-Fan Wu, Jerin Peter, Chip Huyen, Boyuan Chen, Jan Liphardt

SocioGesture: Real-Time and Adaptive Social Gesture Perception for Human-Robot Interaction

Read the original on arXiv Computer Vision →

SocioGesture is a real‑time, adaptive system for recognizing social gestures in human‑robot interaction. It employs a compact, confidence‑aware body‑hand skeleton representation and a lightweight dual‑stream model that fuses body motion with hand articulation, enabling low‑latency onboard recognition. The model is trained with occlusion‑aware skeleton corruption to handle missing hands, occluded arms, and unstable keypoints, and it can expand its gesture vocabulary during deployment by saving uncertain interaction segments for offline labeling.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Computer Vision.

arXiv Computer Vision
Aug 27

Pose-Anchored Optical Flow for Low-Latency Human Action Anticipation in Human-Robot Teaming

Pose-Anchored Optical Flow for Low-Latency Human Action Anticipation in Human-Robot Teaming proposes PoseOFF, a representation that captures local motion around human joints by conditioning optical flow extraction on pose. This structured motion representation aligns with human kinematics and improves early action recognition accuracy across multiple datasets and backbones. PoseOFF achieves comparable or better performance while observing less of the action sequence, making it suitable for real‑time, resource‑constrained robotic systems.

By Lewis de Zoete Grundy, Chris McCarthy, Christopher Fluke
arXiv Computer Vision
Sep 2

RGB-D Video Generation for Improving Human-to-Robot Object Handover Prediction

The paper introduces Hand2Bot, an RGB‑D video dataset designed for human‑to‑robot handover scenarios, capturing body posture and facial expressions amid real‑world noise. It also proposes PassGen, a generative pipeline using stable video diffusion and an Intention‑Aware Temporal Face Encoder to synthesize realistic handover sequences while maintaining hand‑object consistency. A morphology‑based depth editing strategy is employed to replicate realistic sensor noise, and experiments show that training on PassGen yields high intention identification accuracy, low false trigger rates, and robust zero‑shot transfer to a physical robot platform.

By Tianyu Sun, Zhoujie Fu, Zihui Gao, Bang Zhang, Guosheng Lin