arXiv AI By Zeyu He, Shiqi Liu, Ke Chen, Yun Yan, Jinzi Wu, Dianqiao Lei, Sirui Wang, ShuRui Peng, Tao Chen, Zhuo Huang, Yu Wu, Yadong Shao, Zhichao Li, Ke Sun, Yang Guan, Keqiang Li, Shengbo Eben Li

Evaluation Is All You Need for Multi-Modal Autonomous Driving

Read the original on arXiv AI →

The paper introduces iDriveVLA, a multi‑modal planning framework for autonomous driving that addresses a generation‑evaluation asymmetry by improving candidate trajectory spaces and providing a unified, safety‑aware evaluator. It combines a Safety‑aware Scorer for risk estimation with a VLM‑guided Modulator that adapts weighting to the scene, and employs an oracle‑aligned progressive training strategy. On the NAVSIM v1 leaderboard, iDriveVLA achieves a new state‑of‑the‑art PDMS score of 94.95, surpassing human‑expert performance.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
23h ago

Efficient Multi-Modal Planning with Reward-Guided Preference Optimization for Autonomous Driving

Efficient Multi-Modal Planning with Reward-Guided Preference Optimization for Autonomous Driving proposes EMPlan, a hybrid trajectory planning method that combines sparse anchors with an offset refinement module for low-latency, high-accuracy predictions. The approach uses a two-stage training paradigm—pretraining followed by reward-guided fine-tuning—to improve safety without extra inference cost, leveraging rule-based reward signals and unpaired preference supervision. EMPlan is evaluated on the non-reactive NAVSIM benchmark, achieving a favorable balance between planning accuracy and efficiency under real-time constraints.

By Chenglin Chen, Lujia Wang, Xinhu Zheng, Jun Ma, Haoang Li
arXiv AI
Sep 21

HERMES: A Holistic End-to-End Risk-Aware Multimodal Embodied System with Vision-Language Models for Long-Tail Autonomous Driving

HERMES is a holistic end‑to‑end multimodal driving framework that incorporates long‑tail semantic knowledge into trajectory planning for autonomous vehicles. It uses a foundation‑model‑assisted annotation pipeline to build Long‑Tail Scene Context and Long‑Tail Planning Context, capturing hazard‑centric scene information, maneuver intent, and risk‑aware guidance. A Tri‑Modal Driving Module then fuses multi‑view visual observations, historical ego‑motion, and long‑tail semantic instructions to generate intent‑ and risk‑aware trajectories, achieving consistent performance gains on a large‑scale real‑world long‑tail driving benchmark.

By Weizhe Tang, Junwei You, Jiaxi Liu, Zhaoyi Wang, Rui Gan, Zilin Huang, Feng Wei, Bin Ran