arXiv Machine Learning By Jonghoon Lee, Seong Hyeon Park, Byungwoo Jeon, Minha Lee, Jinwoo Shin

Pose6DAug: Physically Plausible Multi-view Object Swapping for Robot Data Augmentation

Read the original on arXiv Machine Learning →

arXiv:2606. 20118v1 Announce Type: cross Abstract: Vision-language-action (VLA) policies have shown strong potential for general-purpose manipulation, yet they often fail on novel, out-of-distribution objects whose appearance or geometry deviates from the training distribution.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv AI
Sep 24

InfiNoVA: Infinite Novel View Augmentation for Viewpoint Invariant Robot Policies

InfiNoVA is a data‑augmentation framework that transforms synchronized multi‑camera demonstrations into a dense, geometrically consistent set of training views by reconstructing each manipulation trajectory as a time‑varying 3D Gaussian. The method renders novel observations from sampled camera poses while preserving the original state‑action pairs, improving frame‑level fidelity and temporal consistency compared to generative synthesis. Across four real‑world manipulation tasks, policies trained with InfiNoVA achieve 5.4× higher average success under unseen randomized viewpoints than VISTA‑based augmentation and 1.7× higher success than training on all five physical camera views.

By Sai Puneeth Reddy Gottam, Elmar Rueckert, Vedant Dave
arXiv Computer Vision
3d ago

BIND: Binding 3D Robot Actions to 2D Image Features

arXiv:2609.38443v1 Announce Type: cross Abstract: We introduce BIND, a new action representation for visuomotor robot policies that binds 3D robot actions to their corresponding 2D image features, yi...

By Cameron Smith, Arsh Tangri, Vitor Guizilini, Yue Wang, Zubair Irshad, Sergey Zakharov
Hugging Face Trending Papers
Jun 1

RoboDream: Compositional World Models for Scalable Robot Data Synthesis

Scaling robot learning requires large-scale, diverse demonstrations, yet real-world data collection via teleoperation remains prohibitively expensive and time-consuming. While video diffusion models offer a promising avenue for data scaling, existing generative approaches are often limited to superficial visual augmentation, or suffer from embodiment hallucinations that yield physically infeasible motions.

arXiv AI
Jun 2

From Human Videos to Robot Manipulation: A Survey on Scalable Vision-Language-Action Learning with Human-Centric Data

arXiv:2606. 00054v1 Announce Type: cross Abstract: Recent progress in generalizable embodied control has been driven by large-scale pretraining of Vision-Language-Action (VLA) models.

By Zhiyuan Feng, Qixiu Li, Huizhi Liang, Rushuai Yang, Yichao Shen, Zhiying Du, Zhaowei Zhang, Yu Deng, Li Zhao, Hao Zhao, Zongqing Lu, Oier Mees, Marc Pollefeys, Jiaolong Yang, Baining Guo
arXiv Machine Learning
Jun 10

Dexterous Point Policy: Learning Point-based Dexterous Hand Policies from Human Demonstrations

arXiv:2606. 10614v1 Announce Type: cross Abstract: Robotic foundation models pre-trained on human demonstration videos have shown promise, but a significant embodiment gap remains when the resulting policies are deployed on real robots.

By Beomjun Kim, Seong Hyeon Park, Seunghoon Sim, Seungjun Moon, Sanghyeok Lee, Jinwoo Shin
arXiv Machine Learning
Jul 8

PoseVLA: Universal Pose Pretraining for Generalizable Vision-Language-Action Policies

arXiv:2602. 19710v3 Announce Type: replace-cross Abstract: Existing Vision-Language-Action (VLA) models often suffer from feature collapse and low training efficiency because they entangle high-level perception with sparse, embodiment-specific action supervision.

By Haitao Lin, Hanyang Yu, Jingshun Huang, He Zhang, Yonggen Ling, Ping Tan, Xiangyang Xue, Yanwei Fu