arXiv Computer Vision

PRISM: Predictive Representation of Interaction Style and Motion for Social Robot Navigation

arXiv AI
Aug 28

Multi-Person Human Motion Forecasting in Complex Scenes

The paper introduces Object-Conditioned Social Diffusion (OCSD), a conditional diffusion model that unifies motion history, multi‑person interactions, and object cues for human motion forecasting in complex scenes. OCSD employs an object‑conditioning mechanism that modulates denoising at each timestep, enabling fine‑grained human‑object reasoning, and a social encoder that captures interactions among all humans. Experiments on the Humans in Kitchens (HiK) and HOI‑M3 benchmarks show state‑of‑the‑art performance, reducing two‑second path error by 31.3% on HiK and 33.2% on HOI‑M3 compared to prior work, while producing more realistic long‑term forecasts.

By Serdar Ozsoy, Lars Doorenbos, Juergen Gall
arXiv AI
Jun 9

HA-VLN 2.0: An Open Benchmark and Leaderboard for Human-Aware Navigation in Discrete and Continuous Environments with Dynamic Multi-Human Interactions

arXiv:2503. 14229v4 Announce Type: replace Abstract: Vision-and-Language Navigation (VLN) has been studied mainly in either discrete or continuous spaces, with little attention to dynamic, crowded environments.

By Yifei Dong, Fengyi Wu, Qi He, Lingdong Kong, Heng Li, Minghan Li, Zebang Cheng, Yuxuan Zhou, Jingdong Sun, Qi Dai, Alexander G Hauptmann, Zhi-Qi Cheng
arXiv AI
Sep 24

Where Should I Join? Robot Group Joining via Language-Guided Goal Prediction

The paper introduces a language‑guided approach for robots to join human groups by predicting socially compliant joining poses. It uses recursive spectral partitioning to generate candidate group subsets, ranks them with a language‑conditioned image–geometry model, and then applies a goal predictor that incorporates human‑formation priors to produce a multimodal energy–orientation map of feasible robot poses. Experiments on various group scenarios show competitive grounding accuracy, superior joining‑pose prediction, and successful real‑robot demonstrations in both static and dynamic settings.

By Zilin Fang, Zishuo Wang, Gim Hee Lee, David Hsu
arXiv AI
Sep 24

Kairos: Grounded Forecasting of Presence and Directional Flow in 4D Scene Graphs

Kairos extends a hierarchical 3D scene graph to a 4D scene graph, storing for each voxel a directional mixture and a presence rate. Spectral predictors forecast both the probability of people being present and the full directional distribution of their motion at any future query time. The model supports conditional queries via pairwise flow dependence and provides calibrated credible intervals that tighten as observations accumulate.

By Iacopo Catalano, Julio A. Placed, Javier Civera, Jorge Pe\~na Queralta