arXiv Machine Learning

Augmentations for Robust and Efficient Imitation Learning in Streamed Video Games

arXiv:2607. 14200v1 Announce Type: new Abstract: Imitation learning is an appealing way to scale game-playing agents to complex 3D environments by training policies to map visual observations to actions from human demonstrations.

arXiv AI
Jun 2

When Does Predictive Inverse Dynamics Outperform Behavior Cloning?

arXiv:2601. 21718v2 Announce Type: replace-cross Abstract: Behavior cloning (BC) is a practical offline imitation learning method, but it often fails when expert demonstrations are limited.

By Lukas Sch\"afer, Pallavi Choudhury, Abdelhak Lemkhenter, Chris Lovett, Somjit Nath, Luis Fran\c{c}a, Matheus Ribeiro Furtado de Mendon\c{c}a, Alex Lamb, Riashat Islam, Siddhartha Sen, John Langford, Katja Hofmann, Sergio Valcarcel Macua
arXiv Computer Vision
Sep 14

PAVXploreRL: Physical-Action-Visual World Model Reinforcement Learning with Action Exploration

PAVXploreRL introduces a reinforcement learning framework that builds on a pretrained latent world model to explicitly optimize Physical Plausibility, Action Adherence, and Visual Fidelity (PAV) objectives. By combining in‑distribution expert trajectories with noise‑driven out‑of‑distribution action exploration, the method avoids reliance on paired video supervision and improves generalization. Experiments demonstrate a 5.6% average performance gain over pretrained baselines and more reliable policy evaluation with reduced overestimation bias.

By Han Wang, Zijun Wang, Shuoshuo Xue, Rui Cao, Fengjiao Chen, Xiaodan Liang, Roy Ka-Wei Lee
arXiv Computer Vision
Sep 7

Persistent Robot World Models: Stabilizing Multi-Step Rollouts via Reinforcement Learning

The paper introduces a reinforcement learning post‑training scheme that trains robot world models on their own autoregressive rollouts, using a contrastive RL objective adapted from diffusion models. It also proposes a training protocol that compares multiple variable‑length futures, a multi‑view visual fidelity reward, and demonstrates state‑of‑the‑art rollout fidelity on the DROID dataset, outperforming baselines on LPIPS, SSIM, and human preference tests.

By Jai Bardhan, Patrik Drozdik, Josef Sivic, Vladimir Petrik
arXiv AI
Jul 7

Multiplayer Interactive World Models with Representation Autoencoders

arXiv:2607. 05352v1 Announce Type: cross Abstract: We introduce the first multiplayer world model for highly dynamic environments governed by complex physical interactions.

By Anthony Hu, V\'aclav Volhejn, Adrien Ramanana Rahary, Chris Mulder, Aditya Makkar, Am\'elie Royer, Manu Orsini, Alyx Liao, Adam Jelley, Eloi Alonso, Florian Laurent, Fredrik Nor\'en, James Swingos, Jan H\"unermann, Kent Rollins, Lucas Hosseini, Matthieu Le Cauchois, Maxim Peter, Pim de Witte, Tim Brown, Vincent Micheli, Moritz B\"ohle, Gabriel de Marmiesse, Viktoriia Sharmanska, Lucia Specia, Michael Black, Patrick P\'erez