arXiv Computer Vision

Lightweight Pedestrian Head-Orientation Recognition Network for Safe Pedestrian-Vehicle Interaction

arXiv Computer Vision
Sep 18

MTF-Net: Multi-Modal Temporal Feature Fusion Network for Pedestrian Intention Prediction

MTF‑Net is a Multi‑Modal Temporal Feature Fusion Network that jointly models kinematic, appearance, and contextual cues for pedestrian intention prediction. It fuses four modalities—bounding‑box dynamics, human pose keypoints, local context, and scene‑level semantics—within a recurrent framework enhanced by gated linear units (GLUs) and an attention‑guided fusion head. Evaluations on the PIE and JAAD benchmarks show that MTF‑Net outperforms recent transformer‑ and graph‑based models, achieving up to 0.95 AUC on PIE and 0.94 AUC on JAAD while maintaining real‑time performance.

By Md Mahfuzur Rahman, Pengzhan Zhou, A. F. M. Abdun Noor, Md Imam Ahasan, Md Mustafizur Rahman, Fang Qu
arXiv Computer Vision
Sep 11

TrajFusionNet+: Transformer-Based Prediction of Pedestrian Crossing Intention via Fusion of Trajectory Representations and Scene Graphs

TrajFusionNet+ is a transformer-based model that predicts pedestrian crossing intention by fusing sequential trajectory data, visual trajectory overlays, and graph-based scene context. It extends the earlier TrajFusionNet with three attention modules—Sequence, Visual, and Graph—to capture temporal, visual, and relational cues. The model outperforms state‑of‑the‑art methods on the PIE and JAAD datasets and shows better generalization under a joint‑training, separate‑evaluation protocol.

By Fran\c{c}ois G. Landry, Moulay A. Akhloufi
arXiv AI
Sep 15

Pedestrian Crossing Intent Classification From Event-Based Vision Using Convolutional Spiking Neural Networks With Temporal Augmentation

The paper presents an end‑to‑end system that converts driving footage into dynamic vision sensor (DVS) event streams, augments training with simulated DVS data, and trains a convolutional spiking neural network (Conv‑SNN) to classify pedestrian crossing intent as crossing or non‑crossing. The Conv‑SNN, trained with a class‑balanced loss and surrogate‑gradient learning, achieves high accuracy and F1 scores on JAAD and CARLA DVS datasets, outperforming or matching prior frame‑based methods while operating on sparse temporal representations. The study details architectural choices, neuron dynamics, and training protocols, and provides a convergence analysis and domain‑transfer evaluation.

By Henok Teklu, Mustafa Sakhai, Maciej Wielgosz, Matej Mertik
arXiv AI
Sep 1

Making the Discrete Continuous: Synthetic RAW Augmentations for Fine-Grained Evaluation of Person Detection Performance in Low Light

arXiv:2605.22455v2 Announce Type: replace-cross Abstract: Real-world deployment of AI vision models is both fueled and limited by the data available for training and testing. Real datasets are sparse...

By Valeria Pais, Malena Mendilaharzu, Daniele Faccio, Luis Oala, Christoph Clausen, Bruno Sanguinetti
arXiv Machine Learning
Sep 22

MMS-VPR: A Fine-Grained Multimodal Street-Level Visual Place Recognition Dataset and Evaluation Benchmark for Dense Pedestrian Environments

arXiv:2505.12254v3 Announce Type: replace-cross Abstract: Existing visual place recognition (VPR) datasets predominantly rely on vehicle-mounted imagery, offer limited multimodal diversity, and under...

By Yiwei Ou, Xiaobin Ren, Ronggui Sun, Guansong Gao, Kaiqi Zhao, Manfredo Manfredini
arXiv Machine Learning
Jun 8

Does Appearance Help? A Systematic Study of Image-Based Re-Identification in Online 3D Multi-Pedestrian Tracking

arXiv:2606. 07233v1 Announce Type: cross Abstract: LiDAR-based 3D Multi-Object Tracking (MOT) typically relies solely on geometric information, which is often insufficient to distinguish between targets during prolonged occlusions or in crowded human-populated environments.

By Eduardo Borges, Lu\'is Garrote, Urbano J. Nunes
arXiv Computer Vision
Sep 4

An Ensemble-Based Self-Taught Learning Approach for Parking Space Classification Under Limited Data

The paper proposes an ensemble-based self‑taught learning framework for parking space classification that uses unsupervised convolutional autoencoders to learn transferable visual representations from unlabeled data. These learned encoders serve as fixed feature extractors for supervised classification with limited annotated samples, and an ensemble of heterogeneous autoencoders with independent classifier heads is employed to enhance robustness and reduce architectural bias. Experiments on PKLot and CNRPark benchmarks demonstrate that this approach significantly lowers annotation requirements while achieving high accuracies (93–96%) under cross‑dataset evaluation protocols.

By Lucas de Oliveira Cunha, Joelton Deonei Gotz, Paulo Lisboa de Almeida, Andre Gustavo Hochuli