arXiv Computer Vision
Sep 3

Beauty is in the AI of the beholder: MLLMs systematically overrate facial attractiveness

Beauty assessments from Multimodal Large Language Models (MLLMs) are increasingly popular, prompting a study comparing 2,513 human ratings to four commercial AI models—Claude, Gemini, GPT, and Grok. The study found that MLLMs consistently rate faces more favorably and with a narrower range than humans, yet they maintain strong correlations with human judgments and accurately track the rank‑ordering of faces. While all models agree strongly with each other, Grok showed the lowest agreement with human ratings, and only face age emerged as a common predictor of attractiveness across humans and MLLMs.

By Santiago Grandas, Juan Sebastian Cely-Acosta, Mohit Mendiratta, Shafee Hassan, Macken Murphy
arXiv Computer Vision
Aug 25

Emotion Intensity Matters: Generating Realistic Expressions in Virtual Humans with CVAEs

This paper introduces a Conditional Variational Autoencoder (CVAEs) approach that generates realistic, controllable emotional facial expressions for virtual humans. Trained on a small dataset of 7,680 samples covering six basic emotions at low and high intensity, the model learns latent representations that preserve key expressive characteristics across intensity levels. The method enables animators to produce emotionally expressive virtual characters without actor performances or manual artistic effort.

By Vitor Miguel Xavier Peres, Lara Volpato, Gabriel Ferri Scnheider, Soraia Raupp Musse
arXiv Computer Vision
3d ago

Multidimensional Observer Model and Perceptual Dimensions of Human Image Quality Assessment

The paper introduces a multidimensional observer model that represents images as distributions in a latent perceptual space and models human image quality judgment as comparisons of noisy samples. By aligning the model with neural representations in the primate ventral stream and fitting it to large-scale behavioral data, the authors demonstrate that the perceptual space required for human quality assessment is extremely low-dimensional relative to the image space. The study reveals that the structure of this perceptual space differs between low-level and high-level quality judgments, indicating that humans construct task-dependent perceptual spaces during visual decision making.

By Sheng Zhao, Weikai Lin, Yuhao Zhu
arXiv Computer Vision
Aug 27

Learning Late, Guiding Early: Timestep-Decoupled Semantic Guidance for Fair Face Generation

The paper introduces Semantic Boundary Predictor (SBP), an inference‑time framework that improves demographic fairness in synthetic face generation by applying a single, one‑shot intervention during reverse denoising. SBP learns linear semantic boundaries from late‑stage latent representations and applies them only at the initial noisy latent, leaving the rest of the diffusion process unchanged. Experiments on CelebA‑HQ show significant reductions in fairness disparity—98% for gender, 95% for binary race, and 15% for four‑class race—while preserving image quality across demographic groups.

By Subir Kumar Parida, Rajbabu Velmurugan, Ketan Kotwal, R. S. Sengar, Swati Hiremath