arXiv AI
1d ago

Generative Interactions: Weaving Multiparty Human Motion with Bilevel Latent Dynamics

The paper introduces BRAID, a hierarchical latent-variable model that generates multi-person human motion by explicitly modeling both group-level interaction dynamics and individual behavior conditioned on evolving group context. It treats social motion generation as a meta-transfer learning problem, learning shared interaction priors across datasets and adapting them to arbitrary context sets of observed people and joints. BRAID supports coherent generation under full, sparse, or partial observations and produces compact social-state vectors useful for downstream embodied-agent systems, with evaluations on social forecasting, tracking, in-filling, and response generation.

By Ojas Shirekar, Yash Surange, Agustinas Ju\v{c}as, Chirag Raman
arXiv AI
Aug 28

Multi-Person Human Motion Forecasting in Complex Scenes

The paper introduces Object-Conditioned Social Diffusion (OCSD), a conditional diffusion model that unifies motion history, multi‑person interactions, and object cues for human motion forecasting in complex scenes. OCSD employs an object‑conditioning mechanism that modulates denoising at each timestep, enabling fine‑grained human‑object reasoning, and a social encoder that captures interactions among all humans. Experiments on the Humans in Kitchens (HiK) and HOI‑M3 benchmarks show state‑of‑the‑art performance, reducing two‑second path error by 31.3% on HiK and 33.2% on HOI‑M3 compared to prior work, while producing more realistic long‑term forecasts.

By Serdar Ozsoy, Lars Doorenbos, Juergen Gall
arXiv AI
Sep 24

Where Should I Join? Robot Group Joining via Language-Guided Goal Prediction

The paper introduces a language‑guided approach for robots to join human groups by predicting socially compliant joining poses. It uses recursive spectral partitioning to generate candidate group subsets, ranks them with a language‑conditioned image–geometry model, and then applies a goal predictor that incorporates human‑formation priors to produce a multimodal energy–orientation map of feasible robot poses. Experiments on various group scenarios show competitive grounding accuracy, superior joining‑pose prediction, and successful real‑robot demonstrations in both static and dynamic settings.

By Zilin Fang, Zishuo Wang, Gim Hee Lee, David Hsu