arXiv AI

Creating Impactful Autonomous Driving Datasets: A Strategic Guide from Research Gap to Benchmark

arXiv:2607. 00710v1 Announce Type: cross Abstract: Well-designed autonomous driving datasets have fundamentally shaped research progress, yet existing literature primarily describes what datasets contain rather than how to strategically design impactful ones.

arXiv Computer Vision
4d ago

A Survey on End-to-End Autonomous Driving Training from the Perspectives of Data, Strategy, and Platform

The paper reviews end‑to‑end autonomous driving (E2E‑AD) training, framing it as a Data‑Strategy‑Platform system. It surveys recent advances in data pipelines, learning paradigms, and training infrastructures, and discusses how these layers interact to influence model performance, robustness, and deployability. The authors highlight current limitations and propose a future vision that prioritizes data value, foundation‑driven generalization, and integrated training‑testing loops for more robust, scalable, and trustworthy autonomous driving systems.

By Chengkai Xu, Yiming Cui, Jiaqi Liu, Yicheng Guo, Cheng Qin, Geyuan Zhang, Xinwei Dong, Shiyu Fang, Peng Hang, Jian Sun
arXiv Machine Learning
3d ago

Invent a Dataset: Measuring dataset generation abilities with zero seed

Invent-A-Dataset is a prompt‑based system that generates realistic, large‑scale datasets from a description, targeting the zero‑data regime where no initial data exists. The authors benchmarked it against five leading model APIs across eight task types and up to 20,000 samples, finding that it outperforms competitors with 17% higher quality and 19% more diverse samples. The diversity advantage grows with dataset size, leading to better downstream training performance and consistently higher rankings for fine‑tuned models.

By Shivalika Singh, Andrija Djurisic, Gbemileke Onilude, Sudip Roy, Sara Hooker
arXiv AI
Jun 10

TaCarla: A comprehensive benchmarking dataset for end-to-end autonomous driving

arXiv:2602. 23499v4 Announce Type: replace-cross Abstract: Collecting a high-quality dataset is a critical task that demands meticulous attention to detail, as overlooking certain aspects can render the entire dataset unusable.

By Tugrul Gorgulu, Atakan Dag, M. Esat Kalfaoglu, Halil Ibrahim Kuru, Baris Can Cam, Halil Ibrahim Ozturk, Ozsel Kilinc
Hugging Face Trending Papers
Aug 18

Plug-and-Play Traffic Element Awareness for End-to-End Autonomous Driving

The paper introduces a plug‑and‑play method that injects traffic‑element signals—such as traffic lights and road signs—into end‑to‑end autonomous driving models with minimal architectural changes. By augmenting several public datasets with comprehensive traffic‑element annotations, the authors evaluate this integration across diverse driving paradigms, consistently improving performance on nuScenes, NAVSIM‑v1, NAVSIM‑v2, and Bench2Drive. The approach achieves a new state‑of‑the‑art result on the challenging NAVSIM‑v2 benchmark, demonstrating the broad utility of traffic‑element awareness.

Hugging Face Trending Papers
Sep 3

Understanding Autonomous Driving Datasets by Describing Differences between Image Subsets in Natural Language

The paper introduces set difference captioning for autonomous driving datasets, aiming to generate natural‑language descriptions of differences between two image subsets. It adapts a two‑stage approach to focus on object‑centric patches, enabling attribution of differences to specific objects or categories. A new benchmark, AD‑Diff Bench, is presented to evaluate this method, especially for sparse, real‑world differences, and the authors provide open‑weight models and code for reproducibility.

arXiv Machine Learning
Sep 4

Understanding Autonomous Driving Datasets by Describing Differences between Image Subsets in Natural Language

The paper introduces set difference captioning for autonomous driving datasets, aiming to generate natural‑language descriptions of differences between two image subsets. It adapts a two‑stage approach to focus on object‑centric patches, allowing attribution of differences to specific objects or categories. A new benchmark, AD‑Diff Bench, is presented to evaluate these methods, especially for sparse, real‑world differences, with open‑weight models to ensure reproducibility.

By Julian Truetsch, Felix Hauser, Christoph Stiller, Frank Bieder
Hugging Face Trending Papers
Aug 3

TALSC: Timeliness-Aware Large-Small VLM Collaboration for Infrastructure-Assisted Autonomous Driving

The deployment of Vision-Language Models (VLMs) in autonomous driving (AD) systems is constrained by on-board computing power, restricting vehicles to small VLMs (SVLMs) with limited perception and reasoning capabilities. Infrastructure-assisted AD alleviates this resource constraint by enabling collaboration with large VLMs (LVLMs) at edge servers.

arXiv Machine Learning
Sep 14

Autonomous Research for Open-Ended Problems: A Case Study on Telecom Ticket Retrieval

The paper investigates how fully autonomous machine‑learning research systems can tackle open‑ended, industry‑grade problems, using telecom ticket retrieval as a case study. It finds that while autonomous research excels at hyperparameter tuning, it lacks human intuition and creativity, yet can achieve about 90% of state‑of‑the‑art performance in a fraction of the time and at modest cost. The authors recommend a hybrid approach where human researchers collaborate with autonomous frameworks for optimal results.

By Junghyun Min, Huseyin Uzunalioglu, Mohamed Trabelsi