RGB-D Video Generation for Improving Human-to-Robot Object Handover Prediction
Read the original on arXiv Computer Vision →The paper introduces Hand2Bot, an RGB‑D video dataset designed for human‑to‑robot handover scenarios, capturing body posture and facial expressions amid real‑world noise. It also proposes PassGen, a generative pipeline using stable video diffusion and an Intention‑Aware Temporal Face Encoder to synthesize realistic handover sequences while maintaining hand‑object consistency. A morphology‑based depth editing strategy is employed to replicate realistic sensor noise, and experiments show that training on PassGen yields high intention identification accuracy, low false trigger rates, and robust zero‑shot transfer to a physical robot platform.
Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Computer Vision.