arXiv Machine Learning

G-MARK: Grounded Multi-Agent Reasoning for Cooperative Driving via Knowledge Graphs

arXiv:2608. 19964v1 Announce Type: new Abstract: Autonomous driving systems must operate under partial observability, where safety-critical objects may be occluded or visible only to neighboring connected vehicles.

arXiv AI
Aug 11

CMU-Drive and V2V-VLA: Cooperative Multi-agent Unified Driving with Reasoning Benchmark and Vehicle-to-Vehicle Vision-Language-Action Models

arXiv:2608. 07621v1 Announce Type: new Abstract: Vision-Language-Action (VLA) models have recently achieved impressive performance for end-to-end autonomous driving, yet existing approaches are primarily designed for an individual single autonomous driving agent with limited support for cooperative perception, reasoning, and planning.

By Hsu-kuang Chiu, Stephen F. Smith
arXiv Computer Vision
4d ago

Reasoning models do not yet follow their reasoning in autonomous driving: The KITScenes LongTail Dataset

The paper introduces the KITScenes LongTail dataset, a curated collection of rare driving scenarios designed to evaluate how well reasoning models in autonomous driving follow their own reasoning. The authors find that many current models frequently diverge between the actions they state in their reasoning chains and the actions they actually execute, a phenomenon they term incoherence. They show that when the reasoning and execution disagree, the reasoning is often correct, and enforcing coherence via a kinematic model can improve motion planning, indicating that coherent action based on stated reasoning is essential for trustworthy autonomous driving.

By Royden Wagner, Omer Sahin Tas, Jaime Villa, Felix Hauser, Yinzhe Shen, Marlon Steiner, Dominik Strutz, Carlos Fernandez, Quentin Delfosse, Christoph Weinhuber, Christian Kinzig, Guillermo S. Gutierrez-Cabello, Hendrik K\"onigshof, Fabian Immel, Richard Schwarzkopf, Nils Alexander Rack, Kevin R\"osch, Kaiwen Wang, Jan-Hendrik Pauls, Martin Lauer, Igor Gilitschenski, Holger Caesar, Christoph Stiller
arXiv Machine Learning
Jul 31

LEDGERMIND: Provenance-Constrained Multimodal Agentic Reasoning with a Structured Evidence Ledger

arXiv:2607. 28374v1 Announce Type: new Abstract: Multimodal agents for visual question answering increasingly operate as multi-step trajectories that interleave perception, retrieval, and reasoning, yet evaluation still largely reduces to final-answer accuracy.

By Enjun Du, Hange Zhou, Chenxu Du, Siyi Liu, Zirong Chen, Ziyu Zheng, Yongqi Zhang
Hugging Face Trending Papers
Jul 30

LEDGERMIND: Provenance-Constrained Multimodal Agentic Reasoning with a Structured Evidence Ledger

Multimodal agents for visual question answering increasingly operate as multi-step trajectories that interleave perception, retrieval, and reasoning, yet evaluation still largely reduces to final-answer accuracy. This aggregate signal cannot tell whether a correct answer was reached through grounded evidence, language priors, or accidental error cancellation.

arXiv Computation and Language
1d ago

Beyond Textual Chain-of-Thought: A Survey on Action-Grounded Reasoning in Autonomous Driving

arXiv:2609.01659v1 Announce Type: cross Abstract: Chain-of-thought (CoT) reasoning powers generative models by eliciting intermediate steps before producing an answer. In autonomous driving, the answ...

By Zhengxu Tang, Xiaozhou Zhang, Guofeng Cui, Ziyu Gong, Zi Wang, Yunfei Shi, Ruifeng Deng, Chengzhi Qi, Ke Chen, Sachin Patil, Tianjun Xiao, Langechuan Liu, Pichao Wang