arXiv:2608. 07751v1 Announce Type: cross Abstract: Safe and efficient robot navigation in crowds requires anticipating pedestrian motion despite uncertain and potentially shifting prediction errors.
By Cheng Guo, Mingzhe Ni, Zheng Liang, Yihu Ling, Yuan Hu, Michele Caprio, Daniele Pucci, Wei Pan
arXiv:2606. 02562v1 Announce Type: cross Abstract: Autonomous robots that interact with people must make safe and efficient decisions under human-induced uncertainty, such as their preferences, goals, competency, and willingness to cooperate.
By Haimin Hu
arXiv:2401.05018v3 Announce Type: replace
Abstract: Human motion prediction is a crucial capability for advanced robotic systems that interact with humans. In facilities with dynamic human-robot coll...
By Sarmad Idrees, Seokman Sohn, Jongeun Choi
arXiv:2608. 13555v1 Announce Type: cross Abstract: Humanoid motion tracking is central to teleoperation and whole-body imitation, yet evaluation often disagrees with what people perceive in videos.
By Dairu Liu, Zekun Qi, Jiayu Zeng, Ruixi Yu, Yu Guan, Yintianrun Zhang, Xuchuan Chen, Sikai Liang, Zekai Li, Chenghuai Lin, Xinqiang Yu, Wenyao Zhang, He Wang, Li Yi
The paper introduces HAP, a Hand-Driven Active Perception framework that predicts future six‑degree‑of‑freedom head motion in egocentric settings by conditioning on observed hand motion and inferred target context. HAP constructs a Predictive Target‑Centric Amodal Occlusion Graph to model current and potential occlusions among candidate objects, fuses this with hand and head motion history, and blends the learned trajectory with a constant‑velocity prior. Experiments on a public dataset and a newly released Bottle RGB‑D dataset demonstrate that HAP outperforms baseline methods in head‑motion prediction, highlighting the importance of hand‑driven intention and dynamic occlusion reasoning.
By Yunji Feng, Junyi Ma, Guanzhong Sun, Chenyang Xu, Hesheng Wang
Uncertainty estimation for Vision-Language-Navigation (VLN) models is a critical task since it can help identify ambiguous and unreliable predictions, enabling agents to make safer navigation decision...
arXiv:2503. 14229v4 Announce Type: replace Abstract: Vision-and-Language Navigation (VLN) has been studied mainly in either discrete or continuous spaces, with little attention to dynamic, crowded environments.
By Yifei Dong, Fengyi Wu, Qi He, Lingdong Kong, Heng Li, Minghan Li, Zebang Cheng, Yuxuan Zhou, Jingdong Sun, Qi Dai, Alexander G Hauptmann, Zhi-Qi Cheng
arXiv:2609.17499v1 Announce Type: cross
Abstract: Uncertainty estimation for Vision-Language-Navigation (VLN) models is a critical task since it can help identify ambiguous and unreliable predictions...
By Vicky Feliren, A. Taufiq Asyhari, Muhamad Risqi U. Saputra
arXiv:2606. 15594v1 Announce Type: cross Abstract: We present SLS^2, a framework for safe feedback motion planning from pixels using robust model predictive control (MPC) in learned latent world models.
By Devesh Nath, Anutam Srinivasan, Haoran Yin, Ruitong Jiang, Jeffrey Fang, Glen Chou
arXiv:2603. 25937v2 Announce Type: replace-cross Abstract: Visual Navigation Models (VNMs) promise generalizable, robot navigation by learning from large-scale visual demonstrations.
By Maeva Guerrier, Karthik Soma, Jana Pavlasek, Giovanni Beltrame
arXiv:2610.01162v1 Announce Type: new
Abstract: Reliable video world models could provide scalable predictive environments for robot learning, planning, and evaluation. However, generated robot video...
By Isaiah Milkey, Som Sagar, Aditya Taparia, Xinyuan Liu, Jiqing Wen, Ransalu Senanayake
arXiv:2606. 06627v1 Announce Type: cross Abstract: Human video datasets used for cotraining robot manipulation policies largely consist of curated demonstrations where motions are orchestrated to resemble robot behavior and 3D hand poses are captured with specialized hardware.
By Richard Li, Aditya Prakash, Andrew Wen, Saurabh Gupta, Yilun Du, Pulkit Agrawal