arXiv AI By Mingyang Xu

Fidelity Preference, Not Demographic Preference: A Pixel-Level Attribute-Sensitivity Audit of Image Aesthetic/Preference Scorers

Read the original on arXiv AI →

The study audits four image aesthetic scorers—LAION-Aesthetics, PickScore, ImageReward, and HPSv2—using pixel‑level interventions on skin tone and body type in both synthetic and real images. It finds that most scorers exhibit a fidelity preference: unaltered images receive the highest scores, while perturbations in either direction are penalized in an inverted‑U pattern, and this effect is largely independent of the skin operator. Synthetic‑only audits are misleading, as the apparent preference for darker skin in synthetic faces reverses or weakens when evaluated on real faces, and cross‑scorer results vary widely, underscoring the need for real‑data, within‑image causal isolation to accurately assess demographic bias.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Jun 4

CounterFace: A Synthetic Face Dataset for Fine-Grained Counterfactual Evaluation of Face Recognition Systems

arXiv:2407. 13922v3 Announce Type: replace-cross Abstract: Face recognition (FR) systems are widely deployed in critical applications, making their reliability and robustness across diverse populations and conditions essential.

By Guruprasad Viswanathan Ramesh, Ashish Hooda, Shimaa Ahmed, Harrison J Rosenberg, Ramya Korlakai Vinayak, Kassem Fawaz
arXiv AI
Sep 3

Disease Burden over Skin Tone: Decomposing the Dermatology-AI Generalization Gap

The study examines why dermatology AI models, largely trained on light‑skinned, cancer‑focused images, perform poorly when applied to diverse patient populations. By comparing a cancer‑trained baseline, two dermatology foundation models, and a general‑purpose vision model on tone‑stratified and disease‑shifted datasets, the authors find that disease‑distribution shift, rather than skin‑tone underrepresentation, is the primary cause of generalization failure. Representation analysis shows that cancer‑specialized features lack transferable structure, while dermatology‑pretrained features maintain stronger clustering, and lightweight adaptation with about ten labeled examples per category can recover most performance.

By Nirajan Kunwor, Sanjaya Poudel, Quoc-Huy Trinh, Jahidul Arafat, Sunil Kumar Gaire
arXiv AI
Jun 3

TASTE: A Designer-Annotated Multi-Dimensional Preference Dataset for AI-Generated Graphic Design

arXiv:2605. 20731v2 Announce Type: replace-cross Abstract: Text-to-image models now generate graphic design at production scale, yet their supervision still comes primarily from photo-style preference datasets with a single overall verdict per comparison.

By Haonan Zhu, Elad Hirsch, Alexandria Minetti, Allison Nulty, Purvanshi Mehta