arXiv Machine Learning

MamMA: A Mamba-Based Pedestrian Trajectory Prediction Algorithm Considering Occupancy Map and Pedestrian Awareness States

MamMA is a pedestrian trajectory prediction algorithm that leverages LiDAR-generated occupancy maps and egocentric vision sensor data. It partitions the occupancy map into patches to extract obstacle features and incorporates pedestrian awareness states, which influence perception and speed. Using a Mamba-based model, MamMA predicts future trajectories and outperforms state‑of‑the‑art methods on multiple benchmark datasets.

arXiv AI
Sep 24

Kairos: Grounded Forecasting of Presence and Directional Flow in 4D Scene Graphs

Kairos extends a hierarchical 3D scene graph to a 4D scene graph, storing for each voxel a directional mixture and a presence rate. Spectral predictors forecast both the probability of people being present and the full directional distribution of their motion at any future query time. The model supports conditional queries via pairwise flow dependence and provides calibrated credible intervals that tighten as observations accumulate.

By Iacopo Catalano, Julio A. Placed, Javier Civera, Jorge Pe\~na Queralta
arXiv Computer Vision
Sep 11

MC-DeTra: Motion-Consistent Joint Object Detection and Socially-Aware Trajectory Forecasting in Bird's-Eye-View Images

MC-DeTra is a reimplementation of the DeTra model that jointly performs object detection and socially-aware trajectory forecasting in bird's-eye-view images. It introduces motion-consistency mechanisms that add supervision from each actor’s past motion, surrounding traffic occupancy, and a consistency constraint aligning predicted heading with motion direction. The added losses are train‑only and inference‑safe, improving dynamic trajectory forecasting on the Waymo Open Dataset while maintaining or enhancing detection accuracy.

By Vladislav Diuzhev, Dmitry Yudin
arXiv Computer Vision
Sep 18

Feeling Terrain Before Crossing: World Models for Off-Road Navigation

The paper introduces Feel‑WM, an off‑road navigation world model that incorporates proprioceptive data to predict both visual scenes and the robot’s physical sensations such as slip, tilt, and shake. By learning a future proprioceptive state and failure risk from the robot’s own experience, the model can evaluate planned trajectories using a separable score that balances goal similarity with predicted failure risk. Experiments on real and simulated off‑road data show that Feel‑WM outperforms visual‑only models in both open‑loop planning and closed‑loop navigation for wheeled and legged robots, and it successfully guides a Husky robot around rough terrain on mountain trails where an end‑to‑end policy fails.

By E-In Son, Dong-Wook Kim, Ji-Hoon Hwang, Kangsun Lee, Jisung Bae, Jung-Taak Kim, Seung-Woo Seo
arXiv Computer Vision
Aug 31

iOSPointMapper: RealTime Pedestrian and Accessibility Mapping with Mobile AI

iOSPointMapper is a mobile app that performs real‑time, privacy‑conscious sidewalk mapping using on‑device semantic segmentation, LiDAR depth estimation, and fused GPS/IMU data on recent iPhones and iPads. It detects and localizes sidewalk‑relevant features such as traffic signs, traffic lights, and poles, and includes a user‑guided annotation interface for validating outputs before submission. The anonymized data is transmitted to the Transportation Data Exchange Initiative (TDEI), where it integrates with broader multimodal transportation datasets, and evaluations show the app’s potential for enhanced pedestrian mapping.

By Himanshu Naidu, Yuxiang Zhang, Sachin Mehta, Anat Caspi
arXiv AI
Sep 10

PV-WM: A Heterogeneous Micro-Macro World Model for Articulated Pedestrian-Vehicle Co-Rollout

PV-WM is a history‑only world model that jointly predicts pedestrian root motion, 15‑joint articulation, and vehicle kinematic states in a synchronized heterogeneous state. It uses recurrent updates to generate pedestrian and vehicle motion chunks, reconstructing vehicle boxes from predicted center, heading, and observed extent, and recomputes pedestrian‑vehicle geometry after each transition. Compared to a one‑shot predictor, PV‑WM reduces Root ADE by 12.7% and MPJPE by 14.8%, and across 824 Waymo contexts it lowers Root ADE by 5.2%, MPJPE by 7.6%, P‑V distance error by 11.9%, and oriented‑box closest‑approach error by 5.8%, while using 57.1% fewer parameters, 96.5% fewer FLOPs, and 25.5% lower p95 latency.

By Haozhuang Chi, Jingsong Liang, Ziying Song, Lei Yang, Shihao Li, Haoruo Zhang, Chen Lv
arXiv AI
Jun 2

DeepIPCv3: Event-Aware Multi-Modal Sensor Fusion for Sudden Pedestrian Crossing Avoidance

arXiv:2606. 01277v1 Announce Type: cross Abstract: Current end-to-end autonomous driving systems predominantly rely on frame-based sensors, which suffer from inherent perception latency and motion blur during highly dynamic encounters, specifically sudden pedestrian crossings.

By Oskar Natan, Andi Dharmawan, Aufaclav Zatu Kusuma Frisky, Jazi Eko Istiyanto, Jun Miura