arXiv:2608.24219v1 Announce Type: new
Abstract: Facial attractiveness has been linked to statistical regularities such as symmetry and averageness, suggesting that beauty may depend on the ease with...
By Francisco M. L\'opez, Jochen Triesch
arXiv:2607. 14631v1 Announce Type: cross Abstract: Inferring apparent personality from facial images is important in social scenarios for embodied agents in human-robot interaction.
By Shuhuan Chen, Xiangyu Zhu, Weisong Zhao, Haichao Shi, Xiao-Yu Zhang, Zhen Lei
arXiv:2508. 03483v3 Announce Type: replace-cross Abstract: While prior research on text-to-image generation has predominantly focused on biases in human depictions, demographic bias in generated objects remains relatively underexplored.
By Dasol Choi, Jihwan Lee, Minjae Lee, Minsuk Kahng
arXiv:2606. 07541v1 Announce Type: cross Abstract: Multimodal large language models (MLLMs) have shown strong performance on objective tasks such as video understanding and reasoning.
By Prabal Shrestha, Bohan Jiang, Haoning Xue, Huan Liu, Xinyi Zhou
arXiv:2608.30210v1 Announce Type: cross
Abstract: AI image generators now create face portraits that are hard to tell from real photographs. Vision-language models (VLMs) are increasingly proposed to...
By Sunwhi Kim (Hwasung Medi-Science University, Dept. of Bio-Healthcare), Sunyul Kim (Yonsei University, Graduate School of Engineering, Dept. of Artificial Intelligence), Meounggun Jo (Hoseo University), Jini Tae (Gwangju Institute of Science and Technology, School of Humanities and Social Sciences)
arXiv:2606. 31704v1 Announce Type: cross Abstract: The deployment of face detection models in real-world applications raises important fairness concerns, as these systems may showcase performance disparities across demographic groups.
By Maxime Moussi, Beno\^it Ronval, Siegfried Nijssen, F\'elicien Schiltz
arXiv:2607. 28936v1 Announce Type: cross Abstract: Facial biometric identification relies on the distinctiveness of user attributes within a high-dimensional embedding space.
By Omid Ahmadieh, Nima Karimian
The paper introduces a benchmarking framework that evaluates open-weight Vision‑Language Models (VLMs) for face recognition by treating explanation quality as a core metric. It defines two key criteria for explanations—relevance, meaning reliance on identity‑stable facial features, and faithfulness, meaning alignment with the visible image content without hallucinations. Using this framework, the authors benchmark several VLM families, jointly assessing face verification accuracy and explanation quality, and find that current models still exhibit shortcomings in their explanations, underscoring the importance of explanation metrics for a complete performance assessment.
By Laurent Colbois, S\'ebastien Marcel
The study evaluates whether multimodal large language models (MLLMs) can produce open‑ended aesthetic critiques comparable to humans. Eight open‑weight MLLMs (7 B–397 B) and GPT‑5.5 were tested on 1,227 r/photocritique posts under various prompts, revealing that reference‑based similarity metrics often misrepresent model performance, while shorter critiques and image‑omission had limited impact. Human judges and annotators found the models’ critiques largely different from human ones, noting that models tend to be overly comprehensive and repetitive rather than selective and specific.
By Sajjad Ghiasvand, Maryam Amirizaniani, Haniyeh Ehsani Oskouie, Mahnoosh Alizadeh, Ramtin Pedarsani
The study compares human and vision‑language model (VLM) responses to cross‑modal association tasks, using identical stimuli (a pseudo‑word and two images) and recording both choices and eye movements. While larger VLMs show some alignment with human choices, their attention patterns correlate poorly with human gaze, performing no better than a simple center‑bias baseline. Fine‑tuning VLMs on human choices improves choice alignment but not attention alignment, and training on human gaze improves attention correlation without affecting choice accuracy.
By Sumin Hong, Katsumi Ibaraki, Renee Shi, David Chiang, Toby Jia-Jun Li
PIVOTSBench is a benchmark designed to assess multimodal large language models’ ability to reason about fine‑grained interpersonal relationships. It is constructed from Social‑IQ 2.0 and YouTube data and evaluates models on predicting bidirectional relationship dimensions grounded in psychology research. The benchmark also includes auxiliary tasks that test models’ capacity to identify and use critical visual cues, and it examines the impact of visual modalities, social role information, and different prediction settings on model performance.
By Shuxiang Zhang, Yiting Yin, Wenxuan Song, Yuhang Wu, Miao Liu
The study audits six vision‑language models (VLMs) to assess whether they consistently encode affective qualities of 3D shapes, using Kansei adjective pairs as affective axes. Across ten ShapeNet categories, models show moderate agreement (mean rank correlation 0.36) that is lower than geometric controls but higher than unrelated adjective pairs, with convergence varying widely by category and axis. The authors demonstrate how this audit informs a UI prototype that selectively exposes Kansei descriptors for generative design interfaces.
By Luca Bux, Thiago Rios, Ingo Scholtes, Stefan Menzel