PAVXploreRL introduces a reinforcement learning framework that builds on a pretrained latent world model to explicitly optimize Physical Plausibility, Action Adherence, and Visual Fidelity (PAV) objectives. By combining in‑distribution expert trajectories with noise‑driven out‑of‑distribution action exploration, the method avoids reliance on paired video supervision and improves generalization. Experiments demonstrate a 5.6% average performance gain over pretrained baselines and more reliable policy evaluation with reduced overestimation bias.
By Han Wang, Zijun Wang, Shuoshuo Xue, Rui Cao, Fengjiao Chen, Xiaodan Liang, Roy Ka-Wei Lee
Squint is a visual Soft Actor Critic algorithm designed to accelerate reinforcement learning for robotics. It combines parallel simulation, a distributional critic, resolution squinting, layer normalization, a tuned update-to-data ratio, and an optimized implementation to reduce wall‑clock training time. On the SO‑101 Task Set, Squint trains policies in as little as 15 minutes on a single RTX 3090 GPU, with most tasks converging in under 6 minutes and successfully transferring to a real SO‑101 robot.
By Abdulaziz Almuzairee, Henrik I. Christensen
The paper introduces a reinforcement learning post‑training scheme that trains robot world models on their own autoregressive rollouts, using a contrastive RL objective adapted from diffusion models. It also proposes a training protocol that compares multiple variable‑length futures, a multi‑view visual fidelity reward, and demonstrates state‑of‑the‑art rollout fidelity on the DROID dataset, outperforming baselines on LPIPS, SSIM, and human preference tests.
By Jai Bardhan, Patrik Drozdik, Josef Sivic, Vladimir Petrik
arXiv:2607. 24112v1 Announce Type: new Abstract: We introduce State Transition Pretraining (STP) as a new scaling axis for GUI agents.
By Xiangyan Liu, Kaixin Li, Haonan Wang, Biao Wu, Meng Fang, Longxu Dou, Chao Du, Michael Qizhe Shieh, Tianyu Pang
arXiv:2609.37250v1 Announce Type: cross
Abstract: World-action models (WAMs) couple future visual-state prediction with action generation. By adapting video generators or image-editing models pretrai...
By Yang Zhang, Jiangyuan Zhao, Chenyou Fan, Jiayu Hu, Xiu Yuan, Chenjia Bai, Xiu Li
arXiv:2609.22083v1 Announce Type: new
Abstract: We present MintAct, a family of vision-language models that unifies UI grounding, multi-step navigation across mobile, desktop, and web, and visual too...
By Mingfei Gao, Rui Tian, Haiming Gang, Bohan Zhai, Le Zhang, Yuanzheng Gong, Di Feng, Ege \"Ozsoy, Kaixin Ma, Vishwesh Kirthivasan, O\u{g}uzhan Fatih Kar, Roman Bachmann, Anders Boesen Lindbo Larsen, Afshin Dehghan