OpenAI Blog

Ingredients for robotics research

We’re releasing eight simulated robotics environments and a Baselines implementation of Hindsight Experience Replay, all developed for our research over the past year. We’ve used these environments to train models which work on physical robots.

OpenAI Blog
May 16, 2017

Robots that learn

We’ve created a robotics system, trained entirely in simulation and deployed on a physical robot, which can learn a new task after seeing it done once.

OpenAI Blog
Oct 19, 2017

Generalizing from simulation

Our latest robotics techniques allow robot controllers, trained entirely in simulation and deployed on physical robots, to react to unplanned changes in the environment as they solve simple tasks. That is, we’ve used these techniques to build closed-loop systems rather than open-loop ones as before.

Hugging Face Trending Papers
Sep 17

Learning and Transferring Closed-Loop Robot Software

Closed‑loop robot policies are difficult to design manually because they require complex observation processing, state management, and branching. This study treats complete closed‑loop implementations as reusable execution experience: a coding agent generates policy code from a few demonstrations, iteratively improves it with simulation feedback, and stores the validated implementations. When applying these archived implementations to new tasks, the agent can generate and refine policies using the stored code, target demonstrations, and execution feedback, ultimately producing a frozen policy that runs without further model calls. Across multiple source and target tasks, iterative optimization of the source implementations significantly boosts success rates, demonstrating the value of execution‑improved software for acquiring new policies.

arXiv AI
Sep 18

Learning and Transferring Closed-Loop Robot Software

The paper investigates whether closed‑loop robot software generated and refined by a coding agent can be reused to acquire policies for new tasks. For each source task, the agent creates policy code from a few demonstrations, iteratively improves it with simulation feedback, and stores the validated implementations. When applied to new tasks, the agent uses these archived implementations, additional demonstrations, and execution feedback to produce a final policy that runs without further model calls, achieving higher success rates than starting from scratch or from unoptimized source code.

By So Kuroki, Yujin Tang