Agent as Policy for Robotic Manipulation
Read the original on arXiv Computation and Language →The Flow has not summarised this story yet — read it at arXiv Computation and Language.
The Flow has not summarised this story yet — read it at arXiv Computation and Language.
arXiv:2607. 11119v1 Announce Type: cross Abstract: Robot manipulation is a complex task that requires visual understanding, physical reasoning, planning, and closed-loop control.
arXiv:2607. 15641v1 Announce Type: cross Abstract: Humans combine reasoning and motor control to solve complex manipulation tasks under diverse constraints.
arXiv:2605. 31286v2 Announce Type: replace-cross Abstract: Real-world household robots require Vision-Language-Action (VLA) foundation models that can acquire reusable manipulation skills across diverse objects, task conditions, and household environments.
RAPID is a system that automatically generates, verifies, and refines robot programs from a single visual human demonstration. It infers a testable task specification, action primitives, and an interactive environment, using an object-centric relational program representation to enable reuse beyond the demonstration. The approach was evaluated in simulation on eight contact-rich manipulation tasks and successfully deployed on a real Franka arm, showing strong generalization across object pose, shape, material, and environment.
arXiv:2503. 22122v2 Announce Type: replace-cross Abstract: Vision-language models (VLMs) have demonstrated remarkable capabilities in robotic planning, particularly for long-horizon tasks that require a holistic understanding of the environment for task decomposition.
arXiv:2607. 04162v1 Announce Type: cross Abstract: Open-ended tabletop manipulation requires agents to not only understand natural language but also adapt to dynamic environments and execution failures.