Hugging Face Trending Papers

Benchmarking CLIP for Zero-Shot Face and Periocular Gender Estimation

The paper evaluates CLIP’s zero‑shot gender estimation on full‑face and periocular images from the Adience dataset. Using image‑text similarity with male/female prompts, CLIP achieves 95.54% accuracy on full faces without task‑specific training. Periocular predictions are initially biased toward males, but threshold alignment improves accuracy to 85.29%, and a linear SVM on CLIP features yields a modest 86.17% accuracy, slightly better than prior Adience results.

arXiv AI
Jun 3

Effect of Demographic Bias on Skin Lesion Classification

arXiv:2606. 03214v1 Announce Type: new Abstract: In this study, we evaluate the performance of skin lesion classification using ResNet-based convolutional models, focusing on the impact of demographic bias in training data, particularly variations in patient sex and age.

By Ralf Raumanns, Gerard Schouten, Veronika Cheplygina, Josien P. W. Pluim
arXiv Computer Vision
Sep 24

Gender Bias in Vision-Language In-Context Learning

The paper investigates how in‑context learning (ICL) in large vision‑language models (LVLMs) can amplify gender bias. Using the VL‑BICLE framework, the authors show that gendered ICL demonstrations shift model bias toward the demonstrated gender, especially in tasks involving gendered language such as image captioning and pronoun prediction. They find that similarity‑based retrieval does not mitigate this bias and that replacing real images with synthetic ones from stable diffusion reduces bias without hurting caption quality.

By Tong Xiang, Noa Garcia, Yuta Nakashima
arXiv Computer Vision
Sep 17

Face-voice Association across LAnguages and Gender (FLAG) 2027 Challenge Evaluation Plan

arXiv:2609.17913v1 Announce Type: new Abstract: Face--voice association models may rely on language or gender cues in the voice rather than on speaker-specific voice characteristics, which can lead t...

By Marta Moscati, Swapnil Khandoker, Muhammad Saad Saeed, Shah Nawaz, Fatima Noor, Rohan Kumar Das, Mubashir Noman, Junaid Mir, Muhammad Haroon Yousaf, Khalid Malik, Markus Schedl
arXiv Computation and Language
Sep 2

When Modality Gap Reduction Fails: Prediction-Level Hubness in CLIP

The paper investigates why reducing the modality gap between image and text representations in CLIP does not always improve zero‑shot classification accuracy. It shows that while average alignment improves, the relative decision margins among classes can shift, leading to a prediction‑level hubness where predictions concentrate on a few classes. Experiments across datasets confirm that accuracy drops correlate with increased prediction concentration for both linear and learning‑based gap corrections.

By Shota Sato, Hajime Kiyama, Tosho Hirasawa, Mamoru Komachi
arXiv Computer Vision
Sep 16

ViD: Vision-Dominant Gender Bias Mitigation for Large Vision-Language Models

The paper introduces ViD, a vision‑dominant gender bias mitigation framework for large vision‑language models. ViD uses causal analysis of attention patterns and dual mechanisms—backdoor adjustment and refined token selection—to suppress bias while preserving reasoning and generation quality. Experiments show a 14.7% reduction in gender bias on FACET and significant improvements on MS COCO image captioning, all without extra training overhead.

By Zhipeng Zhao, Zhaoqiang Wei, Peishun Liu, Youwei Zhao, Ruichun Tang