arXiv AI

Learn to Quantify Social Interaction with Constraints for Pedestrian Walking

arXiv:2606. 17897v1 Announce Type: new Abstract: Long-term human path forecasting in crowds is critical for autonomous moving platforms (like autonomous driving cars and social robots) to avoid collision and make high-quality planning.

arXiv Computer Vision
Sep 11

TrajFusionNet+: Transformer-Based Prediction of Pedestrian Crossing Intention via Fusion of Trajectory Representations and Scene Graphs

TrajFusionNet+ is a transformer-based model that predicts pedestrian crossing intention by fusing sequential trajectory data, visual trajectory overlays, and graph-based scene context. It extends the earlier TrajFusionNet with three attention modules—Sequence, Visual, and Graph—to capture temporal, visual, and relational cues. The model outperforms state‑of‑the‑art methods on the PIE and JAAD datasets and shows better generalization under a joint‑training, separate‑evaluation protocol.

By Fran\c{c}ois G. Landry, Moulay A. Akhloufi
arXiv AI
Sep 24

Where Should I Join? Robot Group Joining via Language-Guided Goal Prediction

The paper introduces a language‑guided approach for robots to join human groups by predicting socially compliant joining poses. It uses recursive spectral partitioning to generate candidate group subsets, ranks them with a language‑conditioned image–geometry model, and then applies a goal predictor that incorporates human‑formation priors to produce a multimodal energy–orientation map of feasible robot poses. Experiments on various group scenarios show competitive grounding accuracy, superior joining‑pose prediction, and successful real‑robot demonstrations in both static and dynamic settings.

By Zilin Fang, Zishuo Wang, Gim Hee Lee, David Hsu
arXiv Computer Vision
Sep 25

A Vision-Language Framework for Measuring Social Life on Sidewalks

The paper introduces a vision‑language framework that extracts social indicators from street‑level imagery, converting panoramic views into sidewalk‑facing sideviews with timestamps. Using a VLM‑based activity detection system, it codes each pedestrian across ten observable dimensions, producing a Social Dwelling Index (SDI) that captures grouping, dwelling, activity diversity, and accessibility flags. Applied to over 100,000 sideviews in New York City, the study finds that pedestrian volume and SDI are only weakly correlated, indicating that high foot traffic does not necessarily equate to intense social activity.

By Liu Liu, Andres Sevtsuk
arXiv AI
Aug 28

Multi-Person Human Motion Forecasting in Complex Scenes

The paper introduces Object-Conditioned Social Diffusion (OCSD), a conditional diffusion model that unifies motion history, multi‑person interactions, and object cues for human motion forecasting in complex scenes. OCSD employs an object‑conditioning mechanism that modulates denoising at each timestep, enabling fine‑grained human‑object reasoning, and a social encoder that captures interactions among all humans. Experiments on the Humans in Kitchens (HiK) and HOI‑M3 benchmarks show state‑of‑the‑art performance, reducing two‑second path error by 31.3% on HiK and 33.2% on HOI‑M3 compared to prior work, while producing more realistic long‑term forecasts.

By Serdar Ozsoy, Lars Doorenbos, Juergen Gall
arXiv AI
Aug 25

MPCFormer: A physics-informed data-driven approach for explainable socially-aware autonomous driving

MPCFormer is a physics‑informed, data‑driven framework that explicitly models multi‑vehicle social interaction dynamics for autonomous driving. It uses a Transformer‑based encoder‑decoder to learn discrete state‑space dynamics from naturalistic data, enabling explainable, human‑like behavior planning within a Model Predictive Control (MPC) framework. In open‑loop NGSIM tests, it achieves the lowest trajectory prediction errors (ADE 0.86 m over 5 s), and in closed‑loop intense interaction scenarios it attains a 94.67 % planning success rate, 15.75 % efficiency gain, and reduces collisions from 21.25 % to 0.5 %.

By Jia Hu, Zhexi Lian, Xuerun Yan, Ruiang Bi, Dou Shen, Yu Ruan, Chunlong Xia, Haoran Wang