Automated pain recognition from facial expression could make continuous welfare assessment practical in sheep, but adoption depends on trust: a stockperson cannot act on a score that arrives without j...
arXiv:2601. 21944v3 Announce Type: replace Abstract: The widespread adoption of deep learning models in computer vision has intensified concerns about interpretability.
By Konstantinos P. Panousis, Diego Marcos
arXiv:2608.15404v2 Announce Type: replace
Abstract: Concept Bottleneck Models (CBMs) are designed to make visual classification interpretable by expressing predictions through human-understandable co...
By Yusuf Meric Karadag, Gulay Oklan, Seref Baris Cagliyan, Umut Ozdemir, Emre Akbas
arXiv:2608. 13267v1 Announce Type: cross Abstract: Existing vision-language model (VLM) benchmarks emphasize perception and reasoning accuracy (how well VLMs describe and reason about what they see in an image), with limited attention to behavioral reliability under uncertainty (how they behave when visual evidence is missing or misleading).
By Paul Osemudiame Oamen, Owusu-Banahene Osei, Ananya Mukherjee, Christian Greisinger, Steffen Eger, Pius Onobhayedo, Wei Zhao
arXiv:2601.14172v4 Announce Type: replace-cross
Abstract: We study neural multi-label classification under severe label imbalance through sentence-level detection of the 19 refined Schwartz human val...
By V\'ictor Yeste, Paolo Rosso
EXPL-FR is a lightweight adapter that aligns a vision‑language model’s image encoder with a frozen face‑recognition (FR) embedding space, enabling the FR model to be explained using semantic attribute prompts without any text training. By mapping 978 attribute prompts across 22 categories into the FR space, the method identifies the most detectable concepts—forming a readable semantic signature that better separates identities than the full vocabulary. The approach is evaluated on four FR backbones and two VLM encoders, providing identity‑level, per‑image, and differential explanations, and demonstrates that prompt‑driven audits can rank FR models by per‑ethnicity error and attribute‑change verification cost without requiring labeled data.
By Guray Ozgur, Mustafa Efe Tamyapar, Naser Damer, Fadi Boutros
arXiv:2606. 15779v1 Announce Type: cross Abstract: Multimodal models can name the action units (AUs) behind a facial emotion, but their AU->emotion rationales are typically plausible rather than faithful: nothing forces the AUs a model invokes to be the AUs that actually drive its prediction.
By Van Thong Huynh, Hong Hai Nguyen, Thuy Pham, Trong Nghia Nguyen, Soo-Hyung Kim
arXiv:2609.06419v1 Announce Type: cross
Abstract: Medical vision-language models (VLMs) require confidence that reflects both answer correctness and patient-specific visual evidence. Recent GRPO-base...
By Yangyang Xie, Ke Hao, Jiaqi Liu, Yun Gu, Xinglin Zhang
arXiv:2607. 20379v1 Announce Type: new Abstract: Natural-language autoencoders score explanations of hidden activations by reconstruction: an explanation is deemed faithful if the activation can be regenerated from it.
By Hiskias Dingeto
arXiv:2604. 27720v2 Announce Type: replace Abstract: Vision-language models (VLMs) are increasingly applied to medical visual question answering (Med-VQA), yet whether they can \emph{localize} the evidence behind their answers---a prerequisite for clinical auditability---is poorly characterized.
By Xupeng Chen, Binbin Shi, Chenqian Le, Qifu Yin, Lang Lin, Haowei Ni, Ran Gong, Panfeng Li
arXiv:2605. 03217v2 Announce Type: replace Abstract: Large language models (LLMs) are increasingly deployed in settings that require nuanced ethical reasoning, yet existing bias evaluations treat model outputs as simply "biased" or "unbiased.
By Yash Aggarwal, Atmika Gorti, Vinija Jain, Aman Chadha, Krishnaprasad Thirunarayan, Manas Gaur
arXiv:2609.16247v1 Announce Type: new
Abstract: Large language models sometimes behave in ways resembling human emotional responses, and recent work has identified internal representations that may e...
By Valen Tagliabue, Leonard Dung, Cameron Berg