Why Do Vision Language Models Struggle To Recognize Human Emotions?
arXiv:2604. 15280v2 Announce Type: replace-cross Abstract: Understanding emotions is a fundamental ability for intelligent systems to be able to interact with humans.
arXiv:2603. 03989v2 Announce Type: replace-cross Abstract: When visual evidence is ambiguous, vision models must decide how to interpret face-like patterns.
arXiv:2604. 15280v2 Announce Type: replace-cross Abstract: Understanding emotions is a fundamental ability for intelligent systems to be able to interact with humans.
arXiv:2603. 06054v2 Announce Type: replace-cross Abstract: The use of Vision-Language Models (VLMs) in automated driving applications is becoming increasingly common, with the aim of leveraging their reasoning and generalisation capabilities to handle long-tail scenarios.
arXiv:2609.00868v1 Announce Type: cross Abstract: Vision-language models are evaluated by aggregate accuracy on multimodal benchmarks, a practice that implicitly assumes the model uses its visual inp...
arXiv:2609.05540v1 Announce Type: cross Abstract: Many medical conditions require diagnosis through detailed, multi-context clinical assessment rather than from visual appearance alone. Despite this,...
The paper examines whether model uncertainty aligns with human disagreement on vision tasks. Using multi‑annotator datasets (FER+ and CIFAR‑10H), the authors find that pretrained models rarely reflect the ambiguity humans perceive, with weak correlations between model confidence and human disagreement. Predictive multiplicity offers only modest improvement, indicating that common uncertainty metrics fail to flag ambiguous cases.
arXiv:2608. 07302v1 Announce Type: cross Abstract: Large Vision-Language Models (LVLMs) often suffer from object hallucination, generating objects that are absent from the image.
arXiv:2609.36563v1 Announce Type: new Abstract: Visual emotion recognition commonly assumes that all evidence required for prediction is contained in the observed image or video. Yet the same visible...
arXiv:2608.24430v1 Announce Type: new Abstract: Responsible deployment of face verification systems requires more than accurate decisions: systems should also provide interpretable and auditable evid...
arXiv:2607. 22745v1 Announce Type: cross Abstract: Rapid advances in image generation are eroding the evidentiary value of visual content in settings where authenticity can affect public safety and personal reputation.
arXiv:2603. 24058v2 Announce Type: replace-cross Abstract: Object hallucination in Large Vision-Language Models (LVLMs) severely compromises their reliability in real-world applications, posing a critical barrier to their deployment in high-stakes scenarios such as autonomous driving and medical image analysis.
EXPL-FR is a lightweight adapter that aligns a vision‑language model’s image encoder with a frozen face‑recognition (FR) embedding space, enabling the FR model to be explained using semantic attribute prompts without any text training. By mapping 978 attribute prompts across 22 categories into the FR space, the method identifies the most detectable concepts—forming a readable semantic signature that better separates identities than the full vocabulary. The approach is evaluated on four FR backbones and two VLM encoders, providing identity‑level, per‑image, and differential explanations, and demonstrates that prompt‑driven audits can rank FR models by per‑ethnicity error and attribute‑change verification cost without requiring labeled data.
arXiv:2608. 13167v1 Announce Type: cross Abstract: When visual evidence is occluded or chaotic, models should abstain.