The paper introduces HuRo, a dataset of 630K robotized episodes derived from diverse human videos, created via a pipeline that aligns observations and actions for robotic use. Experiments on four real‑world manipulation tasks show that scaling robotized pretraining boosts task completion from 51.5% to 80.3% and improves out‑of‑distribution performance under spatial and visual shifts. Ablation studies reveal that visual robotization enhances robustness and that end‑to‑end pretraining with retargeted actions outperforms visual‑only transfer.
By Jinho Jeong, Se June Joo, Jaehyun Kang, Dongyun Kim, Yena Kim, Hanjung Kim, Seon Joo Kim
Despite the recent promise in robot control, video generative models suffer from a domain mismatch due to their primary focus on content creation. For example, their design inherently prioritizes visual fidelity and creativity over computational efficiency and physical realism.
The paper introduces Direction-Scale Decomposition (DSD), an action representation that separates translation and rotation increments into direction and scale components before tokenization. DSD is evaluated with uniform binning and a B-spline tokenizer (BEAST) in both simulation and real-world manipulation tasks, showing improved success rates on LIBERO and SimplerEnv, especially under mixed-dataset training. Real-robot experiments confirm performance gains with and without robotics pretraining, supporting DSD as an effective representation for discrete-token vision-language-action models.
By Yufei Duan, Hang Yin, Alberta Longhini, Chao Tang, Danica Kragic
arXiv:2606. 06627v1 Announce Type: cross Abstract: Human video datasets used for cotraining robot manipulation policies largely consist of curated demonstrations where motions are orchestrated to resemble robot behavior and 3D hand poses are captured with specialized hardware.
By Richard Li, Aditya Prakash, Andrew Wen, Saurabh Gupta, Yilun Du, Pulkit Agrawal
Embodied foundation models are expected to benefit from data scaling like large language models, but face a much tighter data bottleneck. Teleoperated real-robot trajectories remain the dominant pretraining source due to their precise action supervision and embodiment alignment, yet their scalability is limited by high collection cost, acquisition difficulty, and low behavioral and environmental diversity.
arXiv:2607.11498v2 Announce Type: replace-cross
Abstract: Vision-language-action (VLA) models require 3D spatial reasoning, yet RGB observations encode robot-object geometry only implicitly. Lifting...
By Byungkun Lee, Dongyoon Hwang, Dongjin Kim, Hojoon Lee, Hyunseung Kim, Jaegul Choo, Minho Park
Scaling robot learning requires large-scale, diverse demonstrations, yet real-world data collection via teleoperation remains prohibitively expensive and time-consuming. While video diffusion models offer a promising avenue for data scaling, existing generative approaches are often limited to superficial visual augmentation, or suffer from embodiment hallucinations that yield physically infeasible motions.
Ego4WAM investigates how various properties of egocentric human data—such as human‑robot alignment, data duration, task diversity, and supervision type—affect robot learning. The study, conducted under a unified world‑action model framework, shows that aligned demonstrations improve out‑of‑distribution generalization and lower the amount of target‑task robot data needed. It also finds that video‑only supervision remains effective, and that data duration and task diversity influence downstream capabilities in distinct ways, as validated on real robots and RoboDojo.
By Zhihao Sun, Liu Liu, Xinjiang Wang, Haoyi Jiang, Wei Feng, Huiqiang Zhang, Xiaosong Jia, Zhizhong Su, Zuxuan Wu
arXiv:2607.08098v2 Announce Type: replace
Abstract: Event cameras are increasingly adopted in embodied perception for their microsecond temporal resolution, high dynamic range, and resilience to moti...
By Linli Shi, Ruijun Zhang, Ziyun Wang