arXiv Computer Vision By Zhiyun Jiang, Hanyong Wang, Binbin Liang, Yu Xie, Menglong Yang, Wei Li

Absence is Presence: Understanding Visual Scene Negative Events Under Safety Cognitive Constraint

Read the original on arXiv Computer Vision →

The paper introduces a new task called visual scene negative captioning, which aims to describe what should be present in an image but is actually absent, a capability crucial for safety-critical applications. It proposes the CRCD framework, which uses counterfactual reconstruction and contrastive decoding to overcome affirmation bias, limited mental filling, and representation bias. CRCD employs a dual-branch architecture for amodal completion and functional association, along with multi-condition representation learning, to generate accurate negative captions and sets a high-performance baseline for this emerging task.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Computer Vision.

arXiv Computer Vision
Sep 18

Benchmarking MLLMs via Cognitive Expected Scene Graph for Safety-Critical Visual Negation Understanding

The paper introduces a new benchmark for evaluating Multi‑Modal Large Language Models (MLLMs) on visual negation understanding, focusing on safety-critical scenarios. It defines the Scene Negation Understanding under Safety Cognition (SNUS) task and presents a high‑fidelity negative caption dataset that maps dense assertions of localized hazards. The authors also propose the Cognitive Expected Scene Graph (CESG) Score, a polarity‑aware, structure‑grounded metric that remains robust under semantic reversals, revealing that existing models and traditional metrics fail on this task.

By Zhiyun Jiang, Hanyong Wang, Binbin Liang, Yu Xie, Menglong Yang, Wei Li
Hugging Face Trending Papers
Aug 19

When Safety Overrides Vision: Exploring Dynamics between Vision Influence and Safety Alignment in Vision-Language Models

Aligned vision‑language models (VLMs) are designed to combine grounded visual reasoning with safe generation. The study finds that when safety constraints are applied, these models often abstain from answering questions that they could answer under default instruction, yet visual evidence continues to influence the decoding process. The authors show that safety‑induced abstention alters late‑stage hidden‑state dynamics, and that targeted interventions can restore grounded answering without retraining or changing visual inputs.

arXiv AI
Jun 24

UniDrive: A Unified Vision-Language and Grounding Framework for Interpretable Risk Understanding in Autonomous Driving

arXiv:2606. 24759v1 Announce Type: cross Abstract: Recent multimodal large language models (MLLMs) have shown strong potential for autonomous driving scene understanding, yet existing methods still face a fundamental trade-off between temporal reasoning and spatial precision.

By Xiaowei Gao, Pengxiang Li, Yitai Cheng, Ruihan Xu, James Haworth, Stephen Law, Yun Ye
Hugging Face Trending Papers
Jun 23

UniDrive: A Unified Vision-Language and Grounding Framework for Interpretable Risk Understanding in Autonomous Driving

Recent multimodal large language models (MLLMs) have shown strong potential for autonomous driving scene understanding, yet existing methods still face a fundamental trade-off between temporal reasoning and spatial precision. Models that rely on single-frame or low-resolution inputs often miss small, distant, or partially occluded hazards, while language-centric driving models frequently provide limited grounded evidence for their explanations.

arXiv AI
Jun 24

When Language Overwrites Vision: Over-Alignment and Geometric Debiasing in Vision-Language Models

arXiv:2605. 08245v4 Announce Type: replace-cross Abstract: Vision-Language Models (VLMs) increasingly power high-stakes applications, from medical imaging to autonomous systems, yet they routinely hallucinate, confidently describing content not present in the input.

By Harshvardhan Saini, Samyak Jha, Yiming Tang, Dianbo Liu