Variational option discovery algorithms
Read the original on OpenAI Blog →The Flow has not summarised this story yet — read it at OpenAI Blog.
The Flow has not summarised this story yet — read it at OpenAI Blog.
arXiv:2609.36473v1 Announce Type: new Abstract: Temporal abstraction via options can improve exploration in large environments. However, existing option discovery algorithms find subgoals that target...
Wayfarer is a domain‑agnostic, online deep RL agent that discovers options via Laplacian representation learning from high‑dimensional observations and uses them for control. The discovered options improve exploration, accelerate credit assignment, and generalise to unseen settings, leading to faster learning of complex policies. Wayfarer achieves state‑of‑the‑art performance among single‑stream agents on the most challenging Atari 2600 games, especially those requiring long‑horizon exploration such as Montezuma's Revenge and Private Eye.
arXiv:2606. 01655v1 Announce Type: cross Abstract: The Bayesian paradigm offers principled tools for sequential decision-making under uncertainty, but its reliance on a probabilistic model for all parameters can hinder the incorporation of complex structural constraints.
arXiv:2011. 02565v2 Announce Type: replace-cross Abstract: Temporal abstraction allows reinforcement learning agents to represent knowledge and develop strategies over different temporal scales.
arXiv:2606. 02351v1 Announce Type: new Abstract: Bayesian optimization (BO) is a popular and effective approach for tuning expensive, noisy experiments, but requires the formulation of an explicit objective function.
arXiv:2607. 09817v1 Announce Type: new Abstract: We propose a framework for the Markov chain (MC) choice model with panel data, including parameter estimation, personalized choice prediction, and personalized assortment optimization.