Robots that learn
We’ve created a robotics system, trained entirely in simulation and deployed on a physical robot, which can learn a new task after seeing it done once.
We’re releasing eight simulated robotics environments and a Baselines implementation of Hindsight Experience Replay, all developed for our research over the past year. We’ve used these environments to train models which work on physical robots.
We’ve created a robotics system, trained entirely in simulation and deployed on a physical robot, which can learn a new task after seeing it done once.
Our latest robotics techniques allow robot controllers, trained entirely in simulation and deployed on physical robots, to react to unplanned changes in the environment as they solve simple tasks. That is, we’ve used these techniques to build closed-loop systems rather than open-loop ones as before.
Closed‑loop robot policies are difficult to design manually because they require complex observation processing, state management, and branching. This study treats complete closed‑loop implementations as reusable execution experience: a coding agent generates policy code from a few demonstrations, iteratively improves it with simulation feedback, and stores the validated implementations. When applying these archived implementations to new tasks, the agent can generate and refine policies using the stored code, target demonstrations, and execution feedback, ultimately producing a frozen policy that runs without further model calls. Across multiple source and target tasks, iterative optimization of the source implementations significantly boosts success rates, demonstrating the value of execution‑improved software for acquiring new policies.
arXiv:2609.37089v1 Announce Type: new Abstract: Real-world videos provide rich demonstrations of manipulation, but turning them into reusable robot skills requires visually aligned environments, exec...
arXiv:2607. 18488v1 Announce Type: cross Abstract: Reinforcement learning (RL) research has demonstrated success in both physical and simulated domains; however, the predominant methodology remains rooted in simulations.
arXiv:2609.37359v1 Announce Type: cross Abstract: Coding agents can now write, run, and debug programs with little human help. Robot tasks, however, are usually specified by a sentence that leaves ou...
We are releasing Roboschool: open-source software for robot simulation, integrated with OpenAI Gym.
The paper investigates whether closed‑loop robot software generated and refined by a coding agent can be reused to acquire policies for new tasks. For each source task, the agent creates policy code from a few demonstrations, iteratively improves it with simulation feedback, and stores the validated implementations. When applied to new tasks, the agent uses these archived implementations, additional demonstrations, and execution feedback to produce a final policy that runs without further model calls, achieving higher success rates than starting from scratch or from unoptimized source code.
arXiv:2607. 00836v1 Announce Type: cross Abstract: World models are increasingly used in embodied intelligence and generative simulation, yet their scope remains ambiguous across communities.
arXiv:2610.02204v1 Announce Type: cross Abstract: Building reliable robot capabilities across diverse tasks requires substantial human effort to develop and maintain skills, design rewards, and integ...
arXiv:2602. 20220v2 Announce Type: replace-cross Abstract: We investigate what specific design choices enable successful online reinforcement learning (RL) on physical robots.