arXiv AI By Felipe Carlos dos Santos, Eric Antonelo, Gustavo Claudio Karl Couto

Closed-Loop Evaluation of Bird's-Eye-View Maps from Cross-View Transformers as Inputs to Behavior-Cloning Policies

Read the original on arXiv AI →

The paper studies how Bird's‑Eye‑View (BEV) maps predicted by Cross‑View Transformers (CVT) can be used directly as inputs to a Behavior‑Cloning (BC) driving policy in the CARLA simulator. It introduces a six‑channel BEV representation and a Kernel Density Estimation (KDE) weighting scheme to focus learning on underrepresented maneuvers. Closed‑loop tests show that the KDE‑weighted model is the only predicted‑BEV agent to finish an episode without infractions, highlighting that global segmentation scores are poor proxies for driving performance and that prediction quality at critical geometries, especially the route channel, is key to reliable navigation.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Sep 4

LightEMMA: A Longitudinal Evaluation of Vision-Language Models for Autonomous Driving

LightEMMA is a longitudinal evaluation framework that tests the autonomous driving performance of vision‑language models (VLMs) without fine‑tuning or prompt engineering. Using this protocol, the authors evaluated 15 VLMs from five major families on the nuScenes prediction benchmark and found that larger, more capable models do not consistently outperform earlier generations. The study identifies common failure modes such as overreliance on historical actions and difficulty reconciling conflicting visual cues, underscoring the need for domain‑specific adaptation to enhance VLM safety in autonomous driving.

By Zhijie Qiao, Haowei Li, Zhong Cao, Henry X. Liu
arXiv AI
Jun 10

TaCarla: A comprehensive benchmarking dataset for end-to-end autonomous driving

arXiv:2602. 23499v4 Announce Type: replace-cross Abstract: Collecting a high-quality dataset is a critical task that demands meticulous attention to detail, as overlooking certain aspects can render the entire dataset unusable.

By Tugrul Gorgulu, Atakan Dag, M. Esat Kalfaoglu, Halil Ibrahim Kuru, Baris Can Cam, Halil Ibrahim Ozturk, Ozsel Kilinc
arXiv AI
Jul 7

BEVLM: Distilling Semantic Knowledge from LLMs into Bird's-Eye View Representations

arXiv:2603. 06576v2 Announce Type: replace-cross Abstract: The integration of Large Language Models (LLMs) into autonomous driving has attracted growing interest for their strong reasoning and semantic understanding abilities, which are essential for handling complex decision-making and long-tail scenarios.

By Thomas Monninger, Shaoyuan Xie, Qi Alfred Chen, Sihao Ding
arXiv AI
1d ago

Risk-Aware World Modeling with Flow-Guided Occupancy Evolution for Selective Trajectory Planning in Automated Driving

RiskWorld is a risk‑aware world modeling framework that forecasts shared occupancy and selectively replaces planned trajectories in automated driving. It fuses spatial risk fields, temporal actor context, and visual bird’s‑eye‑view features, using flow‑guided evolution to transport occupancy and signed residuals to correct it. In open‑loop planning on nuScenes, RiskWorld achieves the lowest collision rate over a 3‑second horizon and the second‑best average L2 error, running at 11.5 FPS on a single NVIDIA RTX 4090.

By Rongxiang Zeng, Linsen Cai, Jiafu Zhang, Yijie Zhong, Yide Tao, Shuai Wang, Nan Zheng, Hai L. Vu, Alvaro Garcia Hernandez, Yongqi Dong