Currently, autonomous driving object detection models face significant data scarcity and generalization challenges when navigating complex Chinese rural traffic scenarios. To address these limitations, we propose a novel real-synthetic mixed object detection dataset tailored specifically for Chinese rural roads and systematically evaluate the performance of 13 mainstream detectors under different real-to-synthetic data ratios, thereby providing empirical evidence for model selection and data strategy design in rural autonomous driving scenarios.
arXiv:2608.21761v1 Announce Type: new
Abstract: Large collections of street-view imagery provide rich visual information about urban environments, but extracting fine-grained geographic information f...
By Changyu Lee, Yeonsoo Park, Abdullah Alfarrarjeh, Seon Ho Kim
arXiv:2606. 05149v1 Announce Type: cross Abstract: Vehicle body type is a significant determinant of cyclist injury severity in overtaking crashes, yet automated tools for classifying vehicles into injury-risk-relevant categories from naturalistic roadway video do not exist in the open literature.
By Gandhimathi Padmanaban, Fred Feng
Vision-Language Models (VLMs) such as CLIP are now foundational to multimodal systems, yet their robustness to spurious correlations remains poorly understood at scale. We present the first large-scale empirical study of 194 publicly available VLMs, including 16 model families, covering a wide range of model sizes, 24 training datasets, and three evaluation benchmarks, namely ImageNet (overall performance), CelebA (typical single-attribute bias), and UrbanCars (complex multi-attribute biases).
arXiv:2609.00661v1 Announce Type: new
Abstract: Satellite foundation models offer a globally available alternative to census data for commuting origin-destination (OD) generation, yet no study has sy...
By Ashiq Shukoor Iqbal, Wilson Wongso, Flora D. Salim
arXiv:2608.30964v1 Announce Type: new
Abstract: Pretrained vision embeddings are increasingly used as general-purpose representations for modelling how people appraise urban scenes, and are validated...
By Kaizhen Tan, Yuantao Deng
arXiv:2606. 14555v1 Announce Type: cross Abstract: Modern image classifiers widely adopt global average pooling (GAP) followed by a linear classification head.
By Aray Karjauv
arXiv:2606. 31603v1 Announce Type: cross Abstract: Semantic segmentation models struggle with data sparsity and rare or visually diverse regions, e.
By Nikolai R\"ohrich, Julian Glei{\ss}ner, Ahmed H. A. Ibrahim, Silvan Mertes, Tobias Huber
OpenCVL is a large, open dataset for fine-grained cross-view localization, comprising 617,388 ground‑aerial image pairs from 41 European cities. It blends high‑end sensor data with diverse in‑the‑wild images and includes a curation framework to correct pose annotations, enabling reliable evaluation. The dataset also offers cross‑area and snowy test sets to probe generalization, and experiments show that adding noisy in‑the‑wild data improves model performance on clean tests.
By Zimin Xia, Mubariz Zaffar, Junsheng Fu, Alexandre Alahi, Julian F. P. Kooij
arXiv:2506. 16898v2 Announce Type: replace Abstract: Diffusion-based text-to-image models are increasingly used for urban analysis and scenario generation, but their geographic knowledge and representational biases remain poorly understood.
By Ciro Beneduce, Massimiliano Luca, Bruno Lepri
Background-Free Objectness Learning (B-FOR) is a dense, class‑agnostic detection framework that learns objectness without treating unlabeled regions as background. It predicts multi‑scale object‑center and scale fields, using spatially structured soft targets to supervise only reliable annotated areas and introduces displacement‑aware scale fields to model object extent. Experiments on PASCAL VOC, MS‑COCO, and Open Images show B‑FOR improves recall by over +10 AR points compared to prior class‑agnostic baselines, with ablation studies confirming the importance of localized supervision and displacement‑aware scaling.
By Dania Batool, Liliana Lo Presti, Marco La Cascia, Filippo Vella
arXiv:2607. 00090v1 Announce Type: cross Abstract: Urban-scale Visual Place Recognition (VPR) aims to identify the geographic location of a query image by matching it against a geo-tagged database.
By Zhiyao Shu, Jiacheng Yang, Yang Lu, Waishan Qiu, Chuan Li, Da Chen