arXiv AI By Arip Asadulaev, Maksim Bobrin, Salem Lahlou, Dmitry Dylov, Fakhri Karray, Martin Takac

Zero-Shot Off-Policy Learning

Read the original on arXiv AI →

arXiv:2602. 01962v2 Announce Type: replace-cross Abstract: Off-policy learning methods seek to derive an optimal policy directly from a fixed dataset of prior interactions.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Machine Learning
1d ago

Reward as Observation: Learning Reward-Based Policies for Rapid Adaptation

The paper proposes a reward-based policy that relies only on rewards and actions, enabling zero‑shot transfer between source and target environments with entirely different observation spaces. Experiments on Pointmass, Cartpole, 2D Car Racing, and the Stretch robot in Habitat‑Sim show that the policy can adapt to new visual styles or 3D renderings without additional samples. Additionally, the reward policy can guide the training of an observation‑based policy in the target environment.

By Morgan Byrd, Maks Sorokin, Robert Wright, Sehoon Ha