arXiv Computer Vision By Santiago Grandas, Juan Sebastian Cely-Acosta, Mohit Mendiratta, Shafee Hassan, Macken Murphy

Beauty is in the AI of the beholder: MLLMs systematically overrate facial attractiveness

Read the original on arXiv Computer Vision →

Beauty assessments from Multimodal Large Language Models (MLLMs) are increasingly popular, prompting a study comparing 2,513 human ratings to four commercial AI models—Claude, Gemini, GPT, and Grok. The study found that MLLMs consistently rate faces more favorably and with a narrower range than humans, yet they maintain strong correlations with human judgments and accurately track the rank‑ordering of faces. While all models agree strongly with each other, Grok showed the lowest agreement with human ratings, and only face age emerged as a common predictor of attractiveness across humans and MLLMs.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Computer Vision.

arXiv Computer Vision
Sep 1

Frontier vision-language models have overtaken young adults at detecting AI-generated portraits -- but not their calibration

arXiv:2608.30210v1 Announce Type: cross Abstract: AI image generators now create face portraits that are hard to tell from real photographs. Vision-language models (VLMs) are increasingly proposed to...

By Sunwhi Kim (Hwasung Medi-Science University, Dept. of Bio-Healthcare), Sunyul Kim (Yonsei University, Graduate School of Engineering, Dept. of Artificial Intelligence), Meounggun Jo (Hoseo University), Jini Tae (Gwangju Institute of Science and Technology, School of Humanities and Social Sciences)