arXiv Computer Vision

RGB-D Video Generation for Improving Human-to-Robot Object Handover Prediction

The paper introduces Hand2Bot, an RGB‑D video dataset designed for human‑to‑robot handover scenarios, capturing body posture and facial expressions amid real‑world noise. It also proposes PassGen, a generative pipeline using stable video diffusion and an Intention‑Aware Temporal Face Encoder to synthesize realistic handover sequences while maintaining hand‑object consistency. A morphology‑based depth editing strategy is employed to replicate realistic sensor noise, and experiments show that training on PassGen yields high intention identification accuracy, low false trigger rates, and robust zero‑shot transfer to a physical robot platform.

arXiv Computer Vision
Sep 1

ACE-Ego-Hand: Repurposing Video Diffusion Models for Occlusion-Robust Egocentric 3D Hand Motion Recovery

arXiv:2608.20308v2 Announce Type: replace Abstract: Egocentric video offers scalable manipulation data for embodied AI, yet recovering metric 3D hand trajectories remains challenging due to severe ob...

By Yufei Liu, Xixi Wang, Hao Li, Ganlong Zhao, Kaitong Cai, Chengkai Jin, Chunxiao Liu, Jianbo Liu, Siyuan Huang, Xingang Pan, Hongsheng Li
arXiv Computer Vision
Aug 21

DreamHand: Repurposing Video Diffusion Models for Occlusion-Robust Egocentric 3D Hand Motion Recovery

arXiv:2608. 20308v1 Announce Type: new Abstract: Egocentric video offers scalable manipulation data for embodied AI, yet recovering metric 3D hand trajectories remains challenging due to severe object occlusion and frequent out-of-sight gaps.

By Yufei Liu, Xixi Wang, Hao Li, Ganlong Zhao, Kaitong Cai, Chengkai Jin, Chunxiao Liu, Jianbo Liu, Siyuan Huang, Xingang Pan, Hongsheng Li
Hugging Face Trending Papers
Aug 12

HandEdit: A Unified Benchmark for Egocentric Human-to-Robot Dexterous Hand Image Editing

Robotic manipulation with dexterous hands is a cornerstone of Embodied AI, yet its progress is stifled by the high cost of collecting embodiment-aware teleoperation data. While abundant egocentric videos of human hands offer a scalable alternative, the profound discrepancies in appearance, articulation, and camera viewpoints between human and robotic data raise significant challenges for co-training.

arXiv Computer Vision
6d ago

An Evaluation Framework for Generating Multi-View Images of a Person in a Scene

The paper introduces a framework for generating multi‑view images of a person within a natural scene, addressing the scarcity of paired multi‑view datasets for human subjects. It evaluates existing diffusion‑based image‑editing models and finds they often hallucinate head‑turn angles, leading to inconsistent backgrounds. To overcome this, the authors propose the Head Scene Rotation Difference (HSRD) metric, which separates camera movement from head pose changes and enables reliable assessment of 3D spatial parallax for constructing high‑quality synthetic datasets.

By Mahir Majid, Young Kyung Kim, Guillermo Sapiro
arXiv Machine Learning
2d ago

HuRo: Robotizing Human Videos for Scalable VLA Pretraining

The paper introduces HuRo, a dataset of 630K robotized episodes derived from diverse human videos, created via a pipeline that aligns observations and actions for robotic use. Experiments on four real‑world manipulation tasks show that scaling robotized pretraining boosts task completion from 51.5% to 80.3% and improves out‑of‑distribution performance under spatial and visual shifts. Ablation studies reveal that visual robotization enhances robustness and that end‑to‑end pretraining with retargeted actions outperforms visual‑only transfer.

By Jinho Jeong, Se June Joo, Jaehyun Kang, Dongyun Kim, Yena Kim, Hanjung Kim, Seon Joo Kim