arXiv Computer Vision

An Ensemble-Based Self-Taught Learning Approach for Parking Space Classification Under Limited Data

The paper proposes an ensemble-based self‑taught learning framework for parking space classification that uses unsupervised convolutional autoencoders to learn transferable visual representations from unlabeled data. These learned encoders serve as fixed feature extractors for supervised classification with limited annotated samples, and an ensemble of heterogeneous autoencoders with independent classifier heads is employed to enhance robustness and reduce architectural bias. Experiments on PKLot and CNRPark benchmarks demonstrate that this approach significantly lowers annotation requirements while achieving high accuracies (93–96%) under cross‑dataset evaluation protocols.

arXiv AI
Jul 1

Real-Time Source-Free Object Detection

arXiv:2606. 31834v1 Announce Type: cross Abstract: Real-world detectors for autonomous driving, surveillance, and robotics must handle domain-shifts under strict latency and memory constraints, yet existing source-free object detection (SFOD) methods rely on heavyweight architectures that prioritize accuracy alone.

By Sairam VCR, Varun Gopal, Poornima Jain, Vineeth N Balasubramanian, Muhammad Haris Khan
arXiv AI
Sep 10

DGCPath: Distribution-Aware Generative Contrastive Framework for Self-supervised Path Representation Learning -- Extended Version

DGCPath is a Distribution‑Aware Generative Contrastive framework designed for self‑supervised path representation learning. It combines a diffusion‑based view generator, a variational contrastive mechanism that aligns latent features at the distribution level, and a generative cross‑supervision module for view‑level consistency. Experiments on three real‑world trajectory datasets show that DGCPath surpasses state‑of‑the‑art baselines on two downstream tasks, indicating stronger generalization and representation effectiveness.

By Sean Bin Yang, Hao Miao, Zongyi Xu, Jilin Hu, Xiangmeng Wang, Hua Lu, Bin Yang, Christian S. Jensen
arXiv AI
Aug 26

Rethinking Pre-Training and Augmentation for Zero-Shot Cross-City Object Detection

The paper proposes a modular training pipeline for zero‑shot cross‑city object detection that combines a multi‑dataset pre‑training strategy with class‑agnostic objectness distillation and a domain‑resilient augmentation stream featuring a Grayworld transformation. Applied to the RF‑DETR detector, the approach reduces cross‑city distribution gaps while using only 16 GB GPU memory, achieving a 24.29‑point mAP improvement and 1st place on the AI City Challenge Track 6 leaderboard. The authors provide code and data at the referenced GitHub repository.

By Long Hoang Pham, Quoc Pham-Nam Ho, Huy-Hung Nguyen, Duong Nguyen-Ngoc Tran, Ngoc Doan-Minh Huynh, Cu Quoc Le, Hoang-Khang Nguyen, Hyung-Min Jeon, Chi Dai Tran, Son Hong Phan, Duong Khac Vu, Trinh Le Ba Khanh, Jae Wook Jeon
arXiv AI
Jul 28

Multi-model approach for autonomous driving: A comprehensive study on traffic sign-, vehicle- and lane detection and behavioral cloning

arXiv:2603. 09255v2 Announce Type: replace-cross Abstract: Deep learning and computer vision techniques have become increasingly important in the development of self-driving cars.

By Kanishkha Jaisankar, Pranav M. Pawar, Diana Susan Joseph, Raja Muthalagu, Mithun Mukherjee, Dnyaneshawar Mantri, Ramjee Prasad