arXiv Computer Vision

RECAST: From Log Replay to Closed-Loop Driving Simulation with View-Complete Actors

arXiv AI
3d ago

DrivingBench: Can Vision-Language Models Drive a Toyota Corolla?

DrivingBench is the first benchmark that tests general‑purpose vision‑language models on the task of driving a real Toyota Corolla around a parking‑lot cone course. The models receive live camera frames and issue steering and velocity commands, with inference latency counted as part of the challenge. In tests, only GPT‑6 Astra completed the course, while other models showed limited progress or failed to pass half the course.

By Aditya Ramabadran, Simon Mahns, Tobias Gessler