arXiv AI

Towards Zero-Shot Task Transfer with Neurosymbolic World Models

The paper introduces a neurosymbolic world model that separates observation reconstruction from reward prediction, enabling the model to adapt zero‑shot to new reward functions defined over a shared symbolic state space. This approach addresses the task‑dependency of traditional neural world models, which learn latent representations tied to specific training tasks. Experiments show that the neurosymbolic formulation generalises more strongly than purely neural methods.

arXiv Machine Learning
Aug 27

Regularized Latent Dynamics Prediction is a Strong Baseline For Behavioral Foundation Models

The paper introduces Regularized Latent Dynamics Prediction (RLDP), a method that adds orthogonality regularization to self‑supervised next‑state prediction in latent space. RLDP maintains feature diversity, matching or surpassing complex representation learning approaches for zero‑shot reinforcement learning. It also performs robustly in low‑coverage data settings where prior methods fail.

By Pranaya Jajoo, Harshit Sikchi, Siddhant Agarwal, Amy Zhang, Scott Niekum, Martha White
arXiv AI
3d ago

Beyond Prediction: Steering VLM Agents with Retrospective World Modeling

The paper introduces Retrospective World Modeling, a new paradigm for vision‑language‑model (VLM) agents that allows them to reason backward by estimating which action most likely caused a state transition. It proposes the Self‑Consistency Reward (SCR), an intrinsic signal that measures how well a policy action aligns with this retrospective explanation, providing dense transition‑level feedback. Experiments demonstrate that incorporating SCR improves policy robustness and generalization compared to purely prospective world‑modeling approaches.

By Yongjiang Liu, Jie Zhang, Haoyue Zhang, Jingcai Guo, Deze Zeng, Song Guo
arXiv AI
Jun 2

From Demonstrations to Rewards: Test-Time Prompt Optimization for VLM Reward Models

arXiv:2606. 00083v1 Announce Type: cross Abstract: Reinforcement learning relies on accurate reward functions, which are often hand-crafted or even unavailable in real-world applications, such as robotics.

By Christian Gumbsch, Leonardo Barcellona, Lennard Sch\"unemann, Platon Karageorgis, Andrii Zadaianchuk, Zehao Wang, Sergey Zakharov, Fabien Despinoy, Rahaf Aljundi, Efstratios Gavves