arXiv AI

Pix2Act: Image-Space Manipulation Policies with Equivariant Augmentation

arXiv:2607. 11167v1 Announce Type: cross Abstract: Representing manipulation actions as 2D trajectories in the camera plane provides a compact and interpretable basis for learning complex 3D manipulation policies.

arXiv Computer Vision
3d ago

BIND: Binding 3D Robot Actions to 2D Image Features

arXiv:2609.38443v1 Announce Type: cross Abstract: We introduce BIND, a new action representation for visuomotor robot policies that binds 3D robot actions to their corresponding 2D image features, yi...

By Cameron Smith, Arsh Tangri, Vitor Guizilini, Yue Wang, Zubair Irshad, Sergey Zakharov
arXiv Machine Learning
Jun 19

Pose6DAug: Physically Plausible Multi-view Object Swapping for Robot Data Augmentation

arXiv:2606. 20118v1 Announce Type: cross Abstract: Vision-language-action (VLA) policies have shown strong potential for general-purpose manipulation, yet they often fail on novel, out-of-distribution objects whose appearance or geometry deviates from the training distribution.

By Jonghoon Lee, Seong Hyeon Park, Byungwoo Jeon, Minha Lee, Jinwoo Shin
arXiv Computer Vision
Sep 16

GeoLAM: Learning Geometry-Grounded Latent Actions from Unlabeled Human Videos

GeoLAM is a framework that learns geometry‑grounded latent actions from unlabeled human videos. It uses future‑frame reconstruction with a frozen geometric feature hierarchy and motion supervision from a 4D geometry teacher to capture 3D displacement, image‑plane motion, and surface‑orientation changes. After pretraining, the representation serves as transition targets for a world‑action model trained on robot demonstrations, enabling denoised latent actions and executable action chunks without requiring hand‑pose annotations or future‑video generation during deployment.

By Yifan Xie, Hekun Tian, Jinkun Liu, YuAn Wang, Qiao Sun, Wenbo Ding
arXiv AI
Sep 25

KeyGen: Unsupervised Keypoint based Object-Centric Representations for Category-Level Policy Generalization

KeyGen is a framework that learns canonical 3D keypoints from point clouds to create structured, object‑centric representations for policy learning in robotic manipulation. By conditioning a visuomotor diffusion policy on these keypoints and object geometry, it predicts full manipulation trajectories that maintain geometric correspondence across different object instances. Experiments on a photorealistic simulation benchmark with three tasks show that KeyGen outperforms prior methods on both seen and unseen objects, scales with more demonstrations, remains robust to rescaling, and performs well in real‑world manipulation.

By Shuxin Cao, Liquan Wang, Masoud Moghani, Benjamin Joffe, Animesh Garg
arXiv AI
Sep 24

InfiNoVA: Infinite Novel View Augmentation for Viewpoint Invariant Robot Policies

InfiNoVA is a data‑augmentation framework that transforms synchronized multi‑camera demonstrations into a dense, geometrically consistent set of training views by reconstructing each manipulation trajectory as a time‑varying 3D Gaussian. The method renders novel observations from sampled camera poses while preserving the original state‑action pairs, improving frame‑level fidelity and temporal consistency compared to generative synthesis. Across four real‑world manipulation tasks, policies trained with InfiNoVA achieve 5.4× higher average success under unseen randomized viewpoints than VISTA‑based augmentation and 1.7× higher success than training on all five physical camera views.

By Sai Puneeth Reddy Gottam, Elmar Rueckert, Vedant Dave
Hugging Face Trending Papers
Jun 9

ManiSplat: Manipulation Trajectory Synthesis from Monocular Video via Decoupled 3D Gaussian Splatting

Reconstructing dynamic and interactive 3D scenes from real-world observations remains a fundamental challenge in computer vision and robotics. While recent advances in 3D Gaussian Splatting have enabled high-fidelity static reconstruction, extending it to interactive environments with articulated robots and manipulable objects remains difficult due to complex contact interactions and abrupt pose changes.

Hugging Face Trending Papers
Jun 1

RoboDream: Compositional World Models for Scalable Robot Data Synthesis

Scaling robot learning requires large-scale, diverse demonstrations, yet real-world data collection via teleoperation remains prohibitively expensive and time-consuming. While video diffusion models offer a promising avenue for data scaling, existing generative approaches are often limited to superficial visual augmentation, or suffer from embodiment hallucinations that yield physically infeasible motions.