Hugging Face Trending Papers
Jul 18

What Do They See? Interpreting Complex Road Scenarios Through the Eyes of Vision-Language-Action Models for Safe and Trustworthy Autonomous Vehicle Learning

End-to-end autonomous driving models are now able to navigate complex road scenarios, mapping raw sensor observations directly to observed paths for open-loop evaluation and often effective driving in closed-loop evaluation. Yet the internal logic of these safety-critical systems remains largely opaque, due to the complexity of traffic scenes.

arXiv AI
Jul 21

What Do They See? Interpreting Complex Road Scenarios Through the Eyes of Vision-Language-Action Models for Safe and Trustworthy Autonomous Vehicle Learning

arXiv:2607. 16938v1 Announce Type: cross Abstract: End-to-end autonomous driving models are now able to navigate complex road scenarios, mapping raw sensor observations directly to observed paths for open-loop evaluation and often effective driving in closed-loop evaluation.

By Kalpana Panda, Wesley Maia, Vinti Agarwal, Ross Greer
arXiv AI
6d ago

TrafficImag: A Benchmark for Counterfactual Roadside Traffic Video Generation

TrafficImag is the first benchmark designed to evaluate counterfactual roadside traffic video generation, combining a large roadside dataset with an executable protocol that supports behavior reasoning, intervention-aware image editing, and conditional video generation. Each intervention is encoded as an actor-level program specifying target actor, intended behavior, legal route, interaction order, and temporal constraints, allowing a unified evaluation across diverse foundation models. The benchmark assesses four validity dimensions—initial-state correctness, route and behavior validity, interaction consistency, and non-target preservation—and reports that the best models achieve 80.4% macro F1 for reasoning and 55.0% end-to-end success when using a complete condition interface.

By Xiangyu Li, Tianyi Wang, Zhihao Dou, Christian Claudel, Zhaomiao Guo
arXiv AI
Aug 26

Beyond Accuracy: A Dual-Judge Evaluation Protocol for Vision-Language Models in Legally Grounded Tasks

The paper introduces a dual‑judge evaluation protocol for vision‑language models in legally grounded tasks, pairing a 0‑10 quality judge with a strict binary semantic‑equivalence judge. Using a controlled UK traffic‑sign interpretation task, the authors analyze 4,680 evaluations across visibility and occlusion conditions, finding moderate association between judges and an asymmetric Type II error pattern that is most pronounced under heavy occlusion. The protocol requires only one additional LLM call and reveals quality‑trustworthiness signals that single‑judge methods miss.

By Su Myat Noe, Ha Thanh Nguyen, May Myo Zin, Ken Satoh