arXiv Computer Vision By Zhiyun Jiang, Hanyong Wang, Binbin Liang, Yu Xie, Menglong Yang, Wei Li

Benchmarking MLLMs via Cognitive Expected Scene Graph for Safety-Critical Visual Negation Understanding

Read the original on arXiv Computer Vision →

The paper introduces a new benchmark for evaluating Multi‑Modal Large Language Models (MLLMs) on visual negation understanding, focusing on safety-critical scenarios. It defines the Scene Negation Understanding under Safety Cognition (SNUS) task and presents a high‑fidelity negative caption dataset that maps dense assertions of localized hazards. The authors also propose the Cognitive Expected Scene Graph (CESG) Score, a polarity‑aware, structure‑grounded metric that remains robust under semantic reversals, revealing that existing models and traditional metrics fail on this task.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Computer Vision.

arXiv Computer Vision
Sep 18

Absence is Presence: Understanding Visual Scene Negative Events Under Safety Cognitive Constraint

The paper introduces a new task called visual scene negative captioning, which aims to describe what should be present in an image but is actually absent, a capability crucial for safety-critical applications. It proposes the CRCD framework, which uses counterfactual reconstruction and contrastive decoding to overcome affirmation bias, limited mental filling, and representation bias. CRCD employs a dual-branch architecture for amodal completion and functional association, along with multi-condition representation learning, to generate accurate negative captions and sets a high-performance baseline for this emerging task.

By Zhiyun Jiang, Hanyong Wang, Binbin Liang, Yu Xie, Menglong Yang, Wei Li
arXiv AI
Jun 24

UniDrive: A Unified Vision-Language and Grounding Framework for Interpretable Risk Understanding in Autonomous Driving

arXiv:2606. 24759v1 Announce Type: cross Abstract: Recent multimodal large language models (MLLMs) have shown strong potential for autonomous driving scene understanding, yet existing methods still face a fundamental trade-off between temporal reasoning and spatial precision.

By Xiaowei Gao, Pengxiang Li, Yitai Cheng, Ruihan Xu, James Haworth, Stephen Law, Yun Ye
Hugging Face Trending Papers
Jun 23

UniDrive: A Unified Vision-Language and Grounding Framework for Interpretable Risk Understanding in Autonomous Driving

Recent multimodal large language models (MLLMs) have shown strong potential for autonomous driving scene understanding, yet existing methods still face a fundamental trade-off between temporal reasoning and spatial precision. Models that rely on single-frame or low-resolution inputs often miss small, distant, or partially occluded hazards, while language-centric driving models frequently provide limited grounded evidence for their explanations.