arXiv Computer Vision

Robotic Contextual Awareness for Human-Robot Collaboration and Environmental Understanding

The thesis "Robotic Contextual Awareness for Human-Robot Collaboration and Environmental Understanding" addresses the challenge of enabling autonomous mobile robots to operate safely in dynamic, human-centric environments by developing advanced contextual awareness. It presents two complementary research directions: (1) a human re-identification and tracking system that allows a robot to recognize and collaborate with a specific person while ignoring others, and (2) enhanced perceptual capabilities that provide geometric and semantic understanding of the environment for improved motion planning and interaction. These contributions aim to improve robots’ knowledge of their surroundings, facilitating smoother and more natural collaboration with humans.

arXiv AI
Sep 17

HINT-Plan: Human Intention-Aware Robot Task Planning in Context-Rich Environments using Vision Language Models

HINT-Plan is a new method that integrates human intention prediction into robot task planning by using Vision Language Models to infer high‑level human intentions from third‑person images. These intentions are converted into goal states and combined with hierarchical Scene Graphs to formulate joint task‑planning problems in context‑rich environments. In a photorealistic simulation, HINT-Plan achieved a 69.71% success rate, outperforming baselines by up to 35.29% and reducing functional conflicts.

By Yuchen Liu, Luigi Palmieri, Lujun Li, Radu State, Ilche Georgievski, Marco Aiello
arXiv Machine Learning
3d ago

STARS: From Spatiotemporal Dynamics to Social Representations in Human-Robot Interaction

arXiv:2609.40245v2 Announce Type: cross Abstract: Robot navigation in dynamic, human-centered environments requires socially-compliant decisions grounded in robust scene understanding. Recent Vision-...

By Nathan Tsoi, Michael J. Munje, Tejas Oberoi, Rishab Maheshwari, Pengen Zheng, Tanush Chauhan, Peter Stone, Joydeep Biswas
arXiv AI
Aug 11

REMAC: Self-Reflective and Self-Evolving Multi-Agent Collaboration for Long-Horizon Robot Manipulation

arXiv:2503. 22122v2 Announce Type: replace-cross Abstract: Vision-language models (VLMs) have demonstrated remarkable capabilities in robotic planning, particularly for long-horizon tasks that require a holistic understanding of the environment for task decomposition.

By Puzhen Yuan, Angyuan Ma, Yunchao Yao, Huaxiu Yao, Masayoshi Tomizuka, Mingyu Ding
arXiv AI
Jul 14

Think When It Matters: Conditional VLM Reasoning for Social Navigation with RL Policies

arXiv:2607. 10991v1 Announce Type: cross Abstract: As mobile robots become more integrated into everyday human environments, social robot navigation is becoming essential for ensuring human comfort, safety, and trust.

By Ali Ahmadi, Hamed Rahimi, Adrien Jacquet Cretides, Marie Samson, Mahdi Khoramshahi, Mohamed Chetouani