arXiv AI

When Visual Evidence is Ambiguous: Pareidolia as a Diagnostic Probe for Vision Models

arXiv:2603. 03989v2 Announce Type: replace-cross Abstract: When visual evidence is ambiguous, vision models must decide how to interpret face-like patterns.

arXiv AI
Aug 10

Probing Visual Concepts in Lightweight Vision-Language Models for Automated Driving

arXiv:2603. 06054v2 Announce Type: replace-cross Abstract: The use of Vision-Language Models (VLMs) in automated driving applications is becoming increasingly common, with the aim of leveraging their reasoning and generalisation capabilities to handle long-tail scenarios.

By Nikos Theodoridis, Reenu Mohandas, Ganesh Sistu, Anthony Scanlan, Ciar\'an Eising, Tim Brophy
arXiv AI
Sep 10

Knowing When Not to Answer: Abstention and Refusal Reasoning in Vision--Language Models

arXiv:2609.05540v1 Announce Type: cross Abstract: Many medical conditions require diagnosis through detailed, multi-context clinical assessment rather than from visual appearance alone. Despite this,...

By Karan Dua, Amit Agarwal, Hitesh Laxmichand Patel, Hansa Meghwani, Jyotika Singh, Ranjeet Gupta, Graham Horwood, Tao Sheng, Avi Sil, Sujith Ravi, Dan Roth
Hugging Face Trending Papers
6d ago

Does Model Uncertainty Track Human Ambiguity? Evidence from Multi-Annotator Vision Benchmarks

The paper examines whether model uncertainty aligns with human disagreement on vision tasks. Using multi‑annotator datasets (FER+ and CIFAR‑10H), the authors find that pretrained models rarely reflect the ambiguity humans perceive, with weak correlations between model confidence and human disagreement. Predictive multiplicity offers only modest improvement, indicating that common uncertainty metrics fail to flag ambiguous cases.

arXiv AI
Jun 16

Mitigating Object Hallucinations in LVLMs via Attention Imbalance Rectification

arXiv:2603. 24058v2 Announce Type: replace-cross Abstract: Object hallucination in Large Vision-Language Models (LVLMs) severely compromises their reliability in real-world applications, posing a critical barrier to their deployment in high-stakes scenarios such as autonomous driving and medical image analysis.

By Han Sun, Qin Li, Peixin Wang, Min Zhang
arXiv Computer Vision
Aug 25

EXPL-FR: Explaining Face Recognition Models via Vision-Language Alignment

EXPL-FR is a lightweight adapter that aligns a vision‑language model’s image encoder with a frozen face‑recognition (FR) embedding space, enabling the FR model to be explained using semantic attribute prompts without any text training. By mapping 978 attribute prompts across 22 categories into the FR space, the method identifies the most detectable concepts—forming a readable semantic signature that better separates identities than the full vocabulary. The approach is evaluated on four FR backbones and two VLM encoders, providing identity‑level, per‑image, and differential explanations, and demonstrates that prompt‑driven audits can rank FR models by per‑ethnicity error and attribute‑change verification cost without requiring labeled data.

By Guray Ozgur, Mustafa Efe Tamyapar, Naser Damer, Fadi Boutros