arXiv AI By Karan Dua, Amit Agarwal, Hitesh Laxmichand Patel, Hansa Meghwani, Jyotika Singh, Ranjeet Gupta, Graham Horwood, Tao Sheng, Avi Sil, Sujith Ravi, Dan Roth

Knowing When Not to Answer: Abstention and Refusal Reasoning in Vision--Language Models

Read the original on arXiv AI →

The Flow has not summarised this story yet — read it at arXiv AI.

arXiv Machine Learning
Sep 3

FairLens: Benchmarking Fairness in Vision-Language Models for High-Stakes Decision-Making

FairLens is a benchmark and evaluation framework that measures fairness and validity of vision‑language models (VLMs) in high‑stakes domains such as hiring, legal, and healthcare. It uses over 100,000 face‑image and question pairs covering gender, race, and age, and assesses responses through demographic parity, soundness, demographic association, and bias in free‑text generation. The study finds that VLMs often make unwarranted inferences from faces rather than abstaining, especially in legal and healthcare contexts, and that small parity gaps can still hide unsafe treatment across groups.

By Vahid Reza Khazaie, Ahmed Y. Radwan, Shaina Raza
Hugging Face Trending Papers
Aug 19

When Safety Overrides Vision: Exploring Dynamics between Vision Influence and Safety Alignment in Vision-Language Models

Aligned vision‑language models (VLMs) are designed to combine grounded visual reasoning with safe generation. The study finds that when safety constraints are applied, these models often abstain from answering questions that they could answer under default instruction, yet visual evidence continues to influence the decoding process. The authors show that safety‑induced abstention alters late‑stage hidden‑state dynamics, and that targeted interventions can restore grounded answering without retraining or changing visual inputs.

arXiv AI
Sep 7

When Seeing Overrides Knowing: Visual Dominance and Deferral-Based Method for Personalized Safety in VLMs

The paper introduces MPS-Bench, a benchmark of 5,181 scenarios from 584 real-world images across 12 high-risk domains, each paired with a hidden user profile, to evaluate personalized safety in vision‑language models (VLMs). Eight leading VLMs were tested and found to almost always respond directly (86‑99%) without seeking missing context, scoring no higher than 2.6/5 on personalized safety. The authors identify a phenomenon called visual dominance, where visual information enters text representations early and suppresses textual risk signals, and propose PRISM, a lightweight input monitor that predicts when a query should be deferred, achieving 0.978 AUC and outperforming all tested models on the safety‑utility Pareto frontier.

By Edward Sun, Yuchen Wu, Zixian Ma, Eric Hanchen Jiang, Yijia Xiao, Xiaoyuan Yi, Ranjay Krishna, Wei Wang, Jindong Wang, Aylin Caliskan