arXiv AI By Danya Li, Xiang Su, Yan Feng, Rico Krueger

Decoding Pedestrian Crossing Intention from Egocentric Vision via Vision Language Models

Read the original on arXiv AI →

arXiv:2606. 09142v1 Announce Type: cross Abstract: Egocentric vision offers a first-person view of human perception and decision making, yet its potential for traffic-safety prediction remains underexplored.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Computer Vision
Sep 11

TrajFusionNet+: Transformer-Based Prediction of Pedestrian Crossing Intention via Fusion of Trajectory Representations and Scene Graphs

TrajFusionNet+ is a transformer-based model that predicts pedestrian crossing intention by fusing sequential trajectory data, visual trajectory overlays, and graph-based scene context. It extends the earlier TrajFusionNet with three attention modules—Sequence, Visual, and Graph—to capture temporal, visual, and relational cues. The model outperforms state‑of‑the‑art methods on the PIE and JAAD datasets and shows better generalization under a joint‑training, separate‑evaluation protocol.

By Fran\c{c}ois G. Landry, Moulay A. Akhloufi
arXiv AI
Jun 29

EXPLORE-Bench: Egocentric Scene Prediction with Long-Horizon Reasoning

arXiv:2603. 09731v3 Announce Type: replace-cross Abstract: Multimodal large language models (MLLMs) are increasingly considered as a foundation for embodied agents, yet it remains unclear whether they can reliably reason about the long-term physical consequences of actions from an egocentric viewpoint.

By Chengjun Yu, Xuhan Zhu, Chaoqun Du, Pengfei Yu, Wei Zhai, Yang Cao, Zheng-Jun Zha
arXiv Computer Vision
Sep 17

Sim-to-Real Traffic Scene Understanding by Decoupling Semantics from Caption Generation with V-JEPA

The paper presents a decoupled framework for sim-to-real traffic scene understanding, separating semantic fact extraction from caption generation. It uses a frozen V-JEPA encoder for predictive scene representations and a lightweight Llama-based predictor for VQA, followed by a training-free structured refinement that leverages statistical priors, inter-question relationships, and temporal consistency. The refined facts are then fed to Qwen3-VL-8B to produce pedestrian and vehicle descriptions, achieving top performance on the 2026 AI City Challenge Track 2 benchmark with 87.09% VQA accuracy and an overall S2 score of 60.0853.

By Nguyen Hoai Thuong Bui, Thanh Nguyen Vo, Trinh Tra Giang Nguyen, Ha Duc Bui