Learning Social Navigation from Internet Videos in the Policy State Space
Read the original on arXiv Machine Learning →The Flow has not summarised this story yet — read it at arXiv Machine Learning.
The Flow has not summarised this story yet — read it at arXiv Machine Learning.
arXiv:2606. 12603v1 Announce Type: cross Abstract: Autonomous long-horizon sidewalk navigation is essential for micro-mobility applications such as robotic food delivery and assistive electronic wheelchairs.
arXiv:2609.09158v1 Announce Type: cross Abstract: We study the problem of navigating cluttered indoor environments with a humanoid robot. Unlike conventional methods that model navigation as a 2D pat...
World models enable agents to reason about future outcomes and learn policies from their knowledge of state transition, but existing approaches primarily focus on reconstructing future observations or...
The paper introduces a Latent World Model (LWM) for robot navigation that predicts action‑conditioned latent feature compatibility instead of reconstructing future observations. By exploiting the correlation between spatial proximity and latent feature similarity, the model evaluates action consequences directly in latent space and supports counterfactual training using sampled action sequences. The learned world model can supervise policy learning from unlabeled video and further improve policies via reinforcement learning entirely within the model, eliminating the need for action annotations and additional environment interaction.
arXiv:2607. 07357v1 Announce Type: cross Abstract: Effective social robot navigation requires sensitivity to human behavior, often revealed through subtle skeletal cues like gait and orientation.
The paper introduces Planning Diffusion Policy Optimization (PDPO), an offline‑to‑online reinforcement‑learning framework that employs a diffusion policy to produce short‑horizon action chunks for robot crowd navigation. PDPO is pretrained on collision‑avoidance demonstrations and fine‑tuned online with PPO, generating five‑step action sequences applied in a receding‑horizon manner. The authors also identify a benchmark artifact where agents can leave the valid domain without explicit boundary constraints, and they mitigate this by treating boundary violations as collisions, leading to improved success rates over strong baselines.