arXiv:2609.13778v1 Announce Type: cross
Abstract: Human trajectory prediction requires modeling both individual motion patterns and social interactions among agents. Existing methods have made substa...
By Jiaheng Chen, Jiaxing Li, Leixia Wang, Jianan Ju, Tinghe Zhang
The paper introduces BRAID, a hierarchical latent-variable model that generates multi-person human motion by explicitly modeling both group-level interaction dynamics and individual behavior conditioned on evolving group context. It treats social motion generation as a meta-transfer learning problem, learning shared interaction priors across datasets and adapting them to arbitrary context sets of observed people and joints. BRAID supports coherent generation under full, sparse, or partial observations and produces compact social-state vectors useful for downstream embodied-agent systems, with evaluations on social forecasting, tracking, in-filling, and response generation.
By Ojas Shirekar, Yash Surange, Agustinas Ju\v{c}as, Chirag Raman
arXiv:2608. 05673v1 Announce Type: new Abstract: Trajectory prediction has shifted toward structured formulations with explicit social modeling.
By Jiaheng Chen, Jiaxing Li, Tinghe Zhang, Chaopeng Guo
The paper introduces Object-Conditioned Social Diffusion (OCSD), a conditional diffusion model that unifies motion history, multi‑person interactions, and object cues for human motion forecasting in complex scenes. OCSD employs an object‑conditioning mechanism that modulates denoising at each timestep, enabling fine‑grained human‑object reasoning, and a social encoder that captures interactions among all humans. Experiments on the Humans in Kitchens (HiK) and HOI‑M3 benchmarks show state‑of‑the‑art performance, reducing two‑second path error by 31.3% on HiK and 33.2% on HOI‑M3 compared to prior work, while producing more realistic long‑term forecasts.
By Serdar Ozsoy, Lars Doorenbos, Juergen Gall
arXiv:2606. 27539v1 Announce Type: cross Abstract: Social media popularity prediction aims to forecast the future reach or influence of online content from early-stage observations.
By Utkarsh Sahu, Zhisheng Qi, Li Zhu, Yizhao Yang, Jun Li, Ryan Rossi, Yu Wang
The paper introduces a language‑guided approach for robots to join human groups by predicting socially compliant joining poses. It uses recursive spectral partitioning to generate candidate group subsets, ranks them with a language‑conditioned image–geometry model, and then applies a goal predictor that incorporates human‑formation priors to produce a multimodal energy–orientation map of feasible robot poses. Experiments on various group scenarios show competitive grounding accuracy, superior joining‑pose prediction, and successful real‑robot demonstrations in both static and dynamic settings.
By Zilin Fang, Zishuo Wang, Gim Hee Lee, David Hsu
arXiv:2607. 07021v1 Announce Type: new Abstract: Humans continuously coordinate with others in dynamic interactions, often through implicit, hard-to-quantify social norms that act as shared tacit expectations among interacting agents.
By Yi Yang, Siyuan Liu, Xin Gao, Huamu Sun, Chao Liu, Qing Zhou, Bingbing Nie
arXiv:2606. 12657v1 Announce Type: new Abstract: Human mobility data is important for transportation, urban planning, and epidemic control, but large-scale trajectory collection is often costly and privacy-constrained, motivating realistic synthetic trajectory generation.
By Siyu Li, Toan Tran, Lingyi Zhao, Khurram Shafique, Li Xiong
arXiv:2606. 17897v1 Announce Type: new Abstract: Long-term human path forecasting in crowds is critical for autonomous moving platforms (like autonomous driving cars and social robots) to avoid collision and make high-quality planning.
By Xiaodan Shi
arXiv:2609.18125v1 Announce Type: new
Abstract: Humans often observe others before interacting and adjust their behavior accordingly. Robot navigation in crowds, however, often represents pedestrians...
By Bo-Han Chen, Hiromu Taketsugu, Norimichi Ukita
arXiv:2603. 22281v2 Announce Type: replace-cross Abstract: Recent progress in latent world models (e.
By Haichao Zhang, Yijiang Li, Shwai He, Tushar Nagarajan, Mingfei Chen, Jianglin Lu, Ang Li, Yun Fu
SocialReasonBench is a new video‑multiple‑choice QA benchmark designed to test socially grounded reasoning in interactive narrative videos. It uses branching gameplay footage from *Detroit: Become Human*, where player choices create alternative social outcomes that can be verified against the game’s script and flowchart. The benchmark includes seven reasoning dimensions—such as intent recognition, emotional empathy, moral dilemma, counterfactual reasoning, and causal antecedent—and employs a multi‑agent pipeline to curate clips, ground answer labels, and generate theory‑guided questions with diagnostic distractors.
By Zheyu Huang, Zijing Shi, Haozhe Luo, Huadong Tang, Mingyu Liu, Meng Fang, Ling Chen