arXiv AI

GEAR: From Dynamic Encoding to Dynamic Activation in Social Trajectory Prediction

arXiv AI
Aug 28

Multi-Person Human Motion Forecasting in Complex Scenes

The paper introduces Object-Conditioned Social Diffusion (OCSD), a conditional diffusion model that unifies motion history, multi‑person interactions, and object cues for human motion forecasting in complex scenes. OCSD employs an object‑conditioning mechanism that modulates denoising at each timestep, enabling fine‑grained human‑object reasoning, and a social encoder that captures interactions among all humans. Experiments on the Humans in Kitchens (HiK) and HOI‑M3 benchmarks show state‑of‑the‑art performance, reducing two‑second path error by 31.3% on HiK and 33.2% on HOI‑M3 compared to prior work, while producing more realistic long‑term forecasts.

By Serdar Ozsoy, Lars Doorenbos, Juergen Gall
arXiv AI
1d ago

Generative Interactions: Weaving Multiparty Human Motion with Bilevel Latent Dynamics

The paper introduces BRAID, a hierarchical latent-variable model that generates multi-person human motion by explicitly modeling both group-level interaction dynamics and individual behavior conditioned on evolving group context. It treats social motion generation as a meta-transfer learning problem, learning shared interaction priors across datasets and adapting them to arbitrary context sets of observed people and joints. BRAID supports coherent generation under full, sparse, or partial observations and produces compact social-state vectors useful for downstream embodied-agent systems, with evaluations on social forecasting, tracking, in-filling, and response generation.

By Ojas Shirekar, Yash Surange, Agustinas Ju\v{c}as, Chirag Raman
arXiv Computer Vision
Sep 4

Drive-HWM: Hierarchical World Models for Dynamic-Latent Guided Autonomous Driving

Drive‑HWM introduces a hierarchical slow‑fast world modeling framework for autonomous driving. The slow model predicts multi‑step future representations, while the fast model jointly predicts the next frame and immediate action using a lightweight multimodal backbone and an autoregressive expert. Dynamic‑Aware Latents, learned through optical‑flow prediction, explicitly capture motion dynamics, and experiments on NAVSIM v1 and v2 show strong driving performance with validated ablation studies.

By Zhaoxin Fan, Tianbao Zhang, Wenjun Wu, Xiaofeng Wang, Yeying Jin, Jian Zhao, Zheng Zhu, Shuicheng Yan