Robots that learn
We’ve created a robotics system, trained entirely in simulation and deployed on a physical robot, which can learn a new task after seeing it done once.
Our latest robotics techniques allow robot controllers, trained entirely in simulation and deployed on physical robots, to react to unplanned changes in the environment as they solve simple tasks. That is, we’ve used these techniques to build closed-loop systems rather than open-loop ones as before.
We’ve created a robotics system, trained entirely in simulation and deployed on a physical robot, which can learn a new task after seeing it done once.
Closed‑loop robot policies are difficult to design manually because they require complex observation processing, state management, and branching. This study treats complete closed‑loop implementations as reusable execution experience: a coding agent generates policy code from a few demonstrations, iteratively improves it with simulation feedback, and stores the validated implementations. When applying these archived implementations to new tasks, the agent can generate and refine policies using the stored code, target demonstrations, and execution feedback, ultimately producing a frozen policy that runs without further model calls. Across multiple source and target tasks, iterative optimization of the source implementations significantly boosts success rates, demonstrating the value of execution‑improved software for acquiring new policies.
We’re releasing eight simulated robotics environments and a Baselines implementation of Hindsight Experience Replay, all developed for our research over the past year. We’ve used these environments to train models which work on physical robots.
arXiv:2607. 18488v1 Announce Type: cross Abstract: Reinforcement learning (RL) research has demonstrated success in both physical and simulated domains; however, the predominant methodology remains rooted in simulations.
The paper investigates whether closed‑loop robot software generated and refined by a coding agent can be reused to acquire policies for new tasks. For each source task, the agent creates policy code from a few demonstrations, iteratively improves it with simulation feedback, and stores the validated implementations. When applied to new tasks, the agent uses these archived implementations, additional demonstrations, and execution feedback to produce a final policy that runs without further model calls, achieving higher success rates than starting from scratch or from unoptimized source code.
We are releasing Roboschool: open-source software for robot simulation, integrated with OpenAI Gym.
Coding agents powered by large language models (LLMs) have shown remarkable abilities to autonomously reason about and achieve goals in the digital world. However, bringing this success to the physica...
arXiv:2609.38982v1 Announce Type: cross Abstract: Coding agents powered by large language models (LLMs) have shown remarkable abilities to autonomously reason about and achieve goals in the digital w...
arXiv:2503.10118v3 Announce Type: replace-cross Abstract: The sim-to-real gap remains a critical challenge in robotics, hindering the deployment of algorithms trained in simulation to real-world syst...
arXiv:2606. 01478v1 Announce Type: cross Abstract: High-quality, large-scale synthetic data from simulations is becoming a cornerstone for pushing the capabilities of robot algorithms.
arXiv:2608. 11221v1 Announce Type: new Abstract: Cyber-physical systems (CPS) are typically developed by multiple stakeholders who produce artefacts tailored to their specific domains of expertise.