Long-term autonomy in human-populated environments requires anticipating whether and how people will move at times a robot has not yet observed. Existing representations of pedestrian motion face a tr...
arXiv:2606. 20209v1 Announce Type: cross Abstract: Joint spatial and temporal understanding of 3D scenes is a crucial requirement for robots deployed in everyday household environments.
By Francesco Argenziano, Miguel Saavedra-Ruiz, Sacha Morin, Charlie Gauthier, Daniele Nardi, Liam Paull
arXiv:2609.08636v1 Announce Type: cross
Abstract: Egocentric 4D interaction forecasting aims to anticipate both where future interactions will occur in 3D and how the human body will move to realize...
By Qiaohui Chu, Haoyu Zhang, Meng Liu, Haoxiang Shi, Dongmei Jiang, Liqiang Nie
arXiv:2605.25059v4 Announce Type: replace
Abstract: Crucial for autonomous exploration, online 3D occupancy prediction and mapping incrementally construct dense spatial representations on the fly. Em...
By Ruoyu Wang, Yong Liu, Jiahan Li, Sheng Tao, Yuhang Lin, Yukai Ma
PV-WM is a history‑only world model that jointly predicts pedestrian root motion, 15‑joint articulation, and vehicle kinematic states in a synchronized heterogeneous state. It uses recurrent updates to generate pedestrian and vehicle motion chunks, reconstructing vehicle boxes from predicted center, heading, and observed extent, and recomputes pedestrian‑vehicle geometry after each transition. Compared to a one‑shot predictor, PV‑WM reduces Root ADE by 12.7% and MPJPE by 14.8%, and across 824 Waymo contexts it lowers Root ADE by 5.2%, MPJPE by 7.6%, P‑V distance error by 11.9%, and oriented‑box closest‑approach error by 5.8%, while using 57.1% fewer parameters, 96.5% fewer FLOPs, and 25.5% lower p95 latency.
By Haozhuang Chi, Jingsong Liang, Ziying Song, Lei Yang, Shihao Li, Haoruo Zhang, Chen Lv
arXiv:2606. 18824v1 Announce Type: cross Abstract: Pedestrian trajectory prediction from an ego-centric camera is challenging since it depends on complex interactions with vehicles and scene context, as well as the intention of the pedestrian.
By Yuxuan Xie, Nicolas Pugeault, Chongfeng Wei, Hubert P. H. Shum, Edmond S. L. Ho
MamMA is a pedestrian trajectory prediction algorithm that leverages LiDAR-generated occupancy maps and egocentric vision sensor data. It partitions the occupancy map into patches to extract obstacle features and incorporates pedestrian awareness states, which influence perception and speed. Using a Mamba-based model, MamMA predicts future trajectories and outperforms state‑of‑the‑art methods on multiple benchmark datasets.
By Juncen Long, Xiaofeng Jin, Gianluca Bardaro, Simone Mentasti, Matteo Matteucci
arXiv:2606. 03609v1 Announce Type: cross Abstract: Embodied agents that navigate cities rely on world models that predict how their surroundings will change as they move.
By Xuhui Lin, Stephen Law, Nanjiang Chen, Kunyao Li, Tao Yang
Camera-only 4D occupancy forecasting enables autonomous vehicles to predict future 3D semantic scenes solely from historical multi-view images, which is critical for driving safety. Even though current methods have achieved good performance, the strong spatial-temporal modeling between the input multi-view frames is still underexplored, which limits the performance of those methods in future 4D forecasting.
arXiv:2607. 08436v1 Announce Type: cross Abstract: Egocentric human data offers scalable supervision for robot manipulation.
By Baoyu Li, Xinchen Yin, Mengying Lin, Yixin Zhang, Danfei Xu
The paper introduces Feel‑WM, an off‑road navigation world model that incorporates proprioceptive data to predict both visual scenes and the robot’s physical sensations such as slip, tilt, and shake. By learning a future proprioceptive state and failure risk from the robot’s own experience, the model can evaluate planned trajectories using a separable score that balances goal similarity with predicted failure risk. Experiments on real and simulated off‑road data show that Feel‑WM outperforms visual‑only models in both open‑loop planning and closed‑loop navigation for wheeled and legged robots, and it successfully guides a Husky robot around rough terrain on mountain trails where an end‑to‑end policy fails.
By E-In Son, Dong-Wook Kim, Ji-Hoon Hwang, Kangsun Lee, Jisung Bae, Jung-Taak Kim, Seung-Woo Seo
arXiv:2609.37476v1 Announce Type: cross
Abstract: Training robust social-navigation policies requires simulators with diverse scene layouts, terrain, and human motion, but constructing such environme...
By Jiaming Wang, Duc Thang Nguyen, Jizhuo Chen, Volodymyr Shcherbyna, Diwen Liu, Zhengcheng Shen, Harold Soh