arXiv AI By Serdar Ozsoy, Lars Doorenbos, Juergen Gall

Multi-Person Human Motion Forecasting in Complex Scenes

Read the original on arXiv AI →

The paper introduces Object-Conditioned Social Diffusion (OCSD), a conditional diffusion model that unifies motion history, multi‑person interactions, and object cues for human motion forecasting in complex scenes. OCSD employs an object‑conditioning mechanism that modulates denoising at each timestep, enabling fine‑grained human‑object reasoning, and a social encoder that captures interactions among all humans. Experiments on the Humans in Kitchens (HiK) and HOI‑M3 benchmarks show state‑of‑the‑art performance, reducing two‑second path error by 31.3% on HiK and 33.2% on HOI‑M3 compared to prior work, while producing more realistic long‑term forecasts.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Jun 29

EXPLORE-Bench: Egocentric Scene Prediction with Long-Horizon Reasoning

arXiv:2603. 09731v3 Announce Type: replace-cross Abstract: Multimodal large language models (MLLMs) are increasingly considered as a foundation for embodied agents, yet it remains unclear whether they can reliably reason about the long-term physical consequences of actions from an egocentric viewpoint.

By Chengjun Yu, Xuhan Zhu, Chaoqun Du, Pengfei Yu, Wei Zhai, Yang Cao, Zheng-Jun Zha