arXiv Computer Vision

Beauty is in the ELBO of the Beholder: A Variational Account of Processing Fluency in Face Perception

arXiv Computer Vision
Sep 3

Beauty is in the AI of the beholder: MLLMs systematically overrate facial attractiveness

Beauty assessments from Multimodal Large Language Models (MLLMs) are increasingly popular, prompting a study comparing 2,513 human ratings to four commercial AI models—Claude, Gemini, GPT, and Grok. The study found that MLLMs consistently rate faces more favorably and with a narrower range than humans, yet they maintain strong correlations with human judgments and accurately track the rank‑ordering of faces. While all models agree strongly with each other, Grok showed the lowest agreement with human ratings, and only face age emerged as a common predictor of attractiveness across humans and MLLMs.

By Santiago Grandas, Juan Sebastian Cely-Acosta, Mohit Mendiratta, Shafee Hassan, Macken Murphy
arXiv Computer Vision
Aug 25

Emotion Intensity Matters: Generating Realistic Expressions in Virtual Humans with CVAEs

This paper introduces a Conditional Variational Autoencoder (CVAEs) approach that generates realistic, controllable emotional facial expressions for virtual humans. Trained on a small dataset of 7,680 samples covering six basic emotions at low and high intensity, the model learns latent representations that preserve key expressive characteristics across intensity levels. The method enables animators to produce emotionally expressive virtual characters without actor performances or manual artistic effort.

By Vitor Miguel Xavier Peres, Lara Volpato, Gabriel Ferri Scnheider, Soraia Raupp Musse
arXiv Computer Vision
3d ago

Multidimensional Observer Model and Perceptual Dimensions of Human Image Quality Assessment

The paper introduces a multidimensional observer model that represents images as distributions in a latent perceptual space and models human image quality judgment as comparisons of noisy samples. By aligning the model with neural representations in the primate ventral stream and fitting it to large-scale behavioral data, the authors demonstrate that the perceptual space required for human quality assessment is extremely low-dimensional relative to the image space. The study reveals that the structure of this perceptual space differs between low-level and high-level quality judgments, indicating that humans construct task-dependent perceptual spaces during visual decision making.

By Sheng Zhao, Weikai Lin, Yuhao Zhu
arXiv Computer Vision
Aug 27

Learning Late, Guiding Early: Timestep-Decoupled Semantic Guidance for Fair Face Generation

The paper introduces Semantic Boundary Predictor (SBP), an inference‑time framework that improves demographic fairness in synthetic face generation by applying a single, one‑shot intervention during reverse denoising. SBP learns linear semantic boundaries from late‑stage latent representations and applies them only at the initial noisy latent, leaving the rest of the diffusion process unchanged. Experiments on CelebA‑HQ show significant reductions in fairness disparity—98% for gender, 95% for binary race, and 15% for four‑class race—while preserving image quality across demographic groups.

By Subir Kumar Parida, Rajbabu Velmurugan, Ketan Kotwal, R. S. Sengar, Swati Hiremath
arXiv Computer Vision
Sep 1

PIU: Proximity-guided Identity Unlearning in ID-Conditioned Diffusion Models

arXiv:2605.22311v2 Announce Type: replace Abstract: Identity-conditioned diffusion models enable high-quality and identity-consistent face generation, but they also raise severe privacy concerns, as...

By Jose Edgar Hernandez Cancino Estrada, Mauro D\'iaz Lupone, \v{Z}iga Emer\v{s}i\v{c}, Vitomir \v{S}truc, Peter Peer, Darian Toma\v{s}evi\'c
arXiv AI
Aug 26

Fidelity Preference, Not Demographic Preference: A Pixel-Level Attribute-Sensitivity Audit of Image Aesthetic/Preference Scorers

The study audits four image aesthetic scorers—LAION-Aesthetics, PickScore, ImageReward, and HPSv2—using pixel‑level interventions on skin tone and body type in both synthetic and real images. It finds that most scorers exhibit a fidelity preference: unaltered images receive the highest scores, while perturbations in either direction are penalized in an inverted‑U pattern, and this effect is largely independent of the skin operator. Synthetic‑only audits are misleading, as the apparent preference for darker skin in synthetic faces reverses or weakens when evaluated on real faces, and cross‑scorer results vary widely, underscoring the need for real‑data, within‑image causal isolation to accurately assess demographic bias.

By Mingyang Xu
arXiv Computer Vision
Sep 3

Learning to Attract and Repel: Dual Quality Margin Learning for Face Recognition (DQM-Face)

The paper introduces Dual Quality Margin Learning for Face Recognition (DQM‑Face), a framework that combines magnitude‑based and semantic quality estimation to refine attraction and repulsion dynamics during training. By integrating squeeze‑and‑excitation semantic attention with dual margins, the method enhances intra‑class compactness and inter‑class separation, yielding a more discriminative feature geometry. Experiments on challenging benchmarks show that DQM‑Face outperforms state‑of‑the‑art face recognition models and that the learned quality signal aligns well with recognition objectives.

By El Ouanas Belabbaci, Bhavesh Wani, Philipp Terh\"orst
arXiv Machine Learning
Jun 2

How Neural Losses Shape VAE Latents

arXiv:2606. 00635v1 Announce Type: new Abstract: Modern VAEs are rarely trained with the pointwise likelihood implied by the standard $\beta$-VAE objective.

By Giorgio Strano, Luca Cerovaz, Michele Mancusi, Tommaso Mencattini, Emanuele Rodol\`a