arXiv AI

AlignFace: Human-Aligned Face Similarity Metric with Interpretable Concept Relations

arXiv:2608. 14130v1 Announce Type: cross Abstract: Computer vision models for generated facial content, such as face editing and privacy protection, increasingly affect people, requiring similarity metrics that serve as faithful proxies for human perception.

arXiv AI
Sep 21

Benchmarking the Explanatory Quality of Open-Weight Vision-Language Models in Face Recognition

The paper introduces a benchmarking framework that evaluates open-weight Vision‑Language Models (VLMs) for face recognition by treating explanation quality as a core metric. It defines two key criteria for explanations—relevance, meaning reliance on identity‑stable facial features, and faithfulness, meaning alignment with the visible image content without hallucinations. Using this framework, the authors benchmark several VLM families, jointly assessing face verification accuracy and explanation quality, and find that current models still exhibit shortcomings in their explanations, underscoring the importance of explanation metrics for a complete performance assessment.

By Laurent Colbois, S\'ebastien Marcel
arXiv Computer Vision
Aug 25

EXPL-FR: Explaining Face Recognition Models via Vision-Language Alignment

EXPL-FR is a lightweight adapter that aligns a vision‑language model’s image encoder with a frozen face‑recognition (FR) embedding space, enabling the FR model to be explained using semantic attribute prompts without any text training. By mapping 978 attribute prompts across 22 categories into the FR space, the method identifies the most detectable concepts—forming a readable semantic signature that better separates identities than the full vocabulary. The approach is evaluated on four FR backbones and two VLM encoders, providing identity‑level, per‑image, and differential explanations, and demonstrates that prompt‑driven audits can rank FR models by per‑ethnicity error and attribute‑change verification cost without requiring labeled data.

By Guray Ozgur, Mustafa Efe Tamyapar, Naser Damer, Fadi Boutros
arXiv Computer Vision
Sep 25

Long-Tail Adaptive Flow Matching with Explicit Conditional Consistency Guidance for Precise Multimodal Face Synthesis

The paper introduces EC²Face, a multimodal face synthesis framework that enhances semantic alignment by combining Explicit Conditional Consistency Guidance (ECCG) and Long‑Tail Adaptive Flow Matching (LAFM). ECCG enforces pixel‑level consistency between generated faces, textual descriptions, and semantic masks, while a temporal dynamic modulation adjusts supervision strength over diffusion timesteps. LAFM reweights spatial optimization signals according to attribute frequency, improving rare attribute synthesis without adding inference overhead. Experiments demonstrate that EC²Face outperforms baselines, achieving a 29.38% improvement in mask accuracy for rare attributes.

By Yushe Cao, Xuechao Zou, Xing Xi, Dianxi Shi, Chun Yu, Junliang Xing