arXiv:2605.05556v2 Announce Type: replace
Abstract: Artificial neural networks trained on visual tasks develop internal representations resembling those of the primate visual system, a discovery that...
By Yash Mehta, Michael F. Bonner
arXiv:2604. 07282v2 Announce Type: replace-cross Abstract: Automated face recognition has made rapid strides over the past decade due to the unprecedented rise of deep neural network (DNN) models that can be trained for domain-specific tasks.
By Fizza Rubab, Yiying Tong, Arun Ross
The paper introduces a benchmarking framework that evaluates open-weight Vision‑Language Models (VLMs) for face recognition by treating explanation quality as a core metric. It defines two key criteria for explanations—relevance, meaning reliance on identity‑stable facial features, and faithfulness, meaning alignment with the visible image content without hallucinations. Using this framework, the authors benchmark several VLM families, jointly assessing face verification accuracy and explanation quality, and find that current models still exhibit shortcomings in their explanations, underscoring the importance of explanation metrics for a complete performance assessment.
By Laurent Colbois, S\'ebastien Marcel
arXiv:2607. 16214v1 Announce Type: cross Abstract: Image descriptions represented with language models (LMs) predict human brain responses to naturalistic images in high-level visual regions, but the factors driving this predictivity remain unclear.
By Anna Bavaresco, Ina Klari\'c, Raquel Fern\'andez, Marie-Francine Moens
arXiv:2606. 11615v1 Announce Type: cross Abstract: The widespread adoption of face recognition (FR) technologies raises serious privacy concerns, as facial data can be exploited without consent.
By Omid Ahmadieh, Nima Karimian
arXiv:2608.24430v1 Announce Type: new
Abstract: Responsible deployment of face verification systems requires more than accurate decisions: systems should also provide interpretable and auditable evid...
By Ana Estrada-Real, Lydia Alapatt, Christoph Busch, Christian Rathgeb
arXiv:2607. 14932v1 Announce Type: cross Abstract: Synthetic face datasets have become effective enough to train face recognition models with accuracy rivaling that of models trained on real photographs.
By Pawe{\l} Borsukiewicz, Daniele Lunghi, Wendk\^uuni C. Ou\'edraogo, Jacques Klein, Tegawend\'e F. Bissyand\'e
EXPL-FR is a lightweight adapter that aligns a vision‑language model’s image encoder with a frozen face‑recognition (FR) embedding space, enabling the FR model to be explained using semantic attribute prompts without any text training. By mapping 978 attribute prompts across 22 categories into the FR space, the method identifies the most detectable concepts—forming a readable semantic signature that better separates identities than the full vocabulary. The approach is evaluated on four FR backbones and two VLM encoders, providing identity‑level, per‑image, and differential explanations, and demonstrates that prompt‑driven audits can rank FR models by per‑ethnicity error and attribute‑change verification cost without requiring labeled data.
By Guray Ozgur, Mustafa Efe Tamyapar, Naser Damer, Fadi Boutros
arXiv:2609.00411v1 Announce Type: new
Abstract: Modern face recognition (FR) owes much of its success to deep neural networks that learn to extract compact identity embeddings from face images. These...
By Fizza Rubab, Yiying Tong, Arun Ross
arXiv:2603.06399v2 Announce Type: replace
Abstract: Facial attribute classification relies on large-scale annotated datasets in which many traits, such as age and expression, are inherently ambiguous...
By Basudha Pal, Zhaoyang Wang, Rama Chellappa
The paper introduces EC²Face, a multimodal face synthesis framework that enhances semantic alignment by combining Explicit Conditional Consistency Guidance (ECCG) and Long‑Tail Adaptive Flow Matching (LAFM). ECCG enforces pixel‑level consistency between generated faces, textual descriptions, and semantic masks, while a temporal dynamic modulation adjusts supervision strength over diffusion timesteps. LAFM reweights spatial optimization signals according to attribute frequency, improving rare attribute synthesis without adding inference overhead. Experiments demonstrate that EC²Face outperforms baselines, achieving a 29.38% improvement in mask accuracy for rare attributes.
By Yushe Cao, Xuechao Zou, Xing Xi, Dianxi Shi, Chun Yu, Junliang Xing
arXiv:2510.01030v2 Announce Type: replace
Abstract: The human ability to translate diverse perceptual and linguistic inputs into structured behavior has been thought to rest on learning robust repres...
By Zach Studdiford, Timothy T. Rogers, Kushin Mukherjee, Siddharth Suresh