arXiv AI

Knowing You at First Glance: Inferring Apparent Personality from Faces

arXiv:2607. 14631v1 Announce Type: cross Abstract: Inferring apparent personality from facial images is important in social scenarios for embodied agents in human-robot interaction.

arXiv AI
Jun 10

Modeling Complex Behaviors: Multi-Personality Composition and Dynamic Switching in Vision-Language Models

arXiv:2606. 11074v1 Announce Type: cross Abstract: With the widespread deployment of Multimodal Large Language Models (MLLMs) in social interaction, understanding and controlling their behavior under complex personality conditions is essential.

By Peiqi Jia (Xi'an Jiaotong University), Haonan Jia (Beihang University), Ziqi Miao (Beihang University), Linkang Du (Xi'an Jiaotong University), Yuntao Wang (Xi'an Jiaotong University), Zhou Su (Xi'an Jiaotong University)
arXiv Computer Vision
Sep 21

Traits Run Deeper: Trait-Specific Asymmetric Fusion for Multimodal Personality Assessment

Traits Run Deeper introduces a personality assessment framework that tailors multimodal fusion to each trait dimension. It comprises a Multimodal Foundation Representation module that uses psychology-informed semantic templates, a Trait-Specific Modality Fusion module that asymmetrically fuses modalities to reduce cross‑modal interference, and a Distribution‑Calibrated Personality Regression module that corrects label imbalance. The approach achieves a ~25% reduction in mean squared error on the AVI Challenge 2026 validation set and wins the Personality Assessment Track.

By Jia Li, Qian Chen, Wei Wang, Xinyu Li, Zhenzhen Hu, Dongsheng Shao, Richang Hong, Meng Wang
arXiv Computer Vision
Sep 3

Beauty is in the AI of the beholder: MLLMs systematically overrate facial attractiveness

Beauty assessments from Multimodal Large Language Models (MLLMs) are increasingly popular, prompting a study comparing 2,513 human ratings to four commercial AI models—Claude, Gemini, GPT, and Grok. The study found that MLLMs consistently rate faces more favorably and with a narrower range than humans, yet they maintain strong correlations with human judgments and accurately track the rank‑ordering of faces. While all models agree strongly with each other, Grok showed the lowest agreement with human ratings, and only face age emerged as a common predictor of attractiveness across humans and MLLMs.

By Santiago Grandas, Juan Sebastian Cely-Acosta, Mohit Mendiratta, Shafee Hassan, Macken Murphy
arXiv Computation and Language
Sep 25

Persona Prompting in Multimodal Urban Perception: Descriptive Convergence and Interpretive Variation

The paper investigates how persona prompting influences the language produced by two multimodal large language models—Qwen3‑VL and Gemma4—when describing urban scenes. Outputs are categorized into descriptive grounding (captions), perception tags, and interpretive framing (justifications). Results show captions largely converge across persona profiles, while justifications differ markedly, especially along economic status, political orientation, and personality dimensions, with economic status producing the greatest variation. Perception tags also reflect attribute similarities, and exploratory topic analysis indicates persona‑specific evaluative emphasis. Overall, persona prompting has a stronger effect on interpretive framing than on descriptive grounding.

By Neemias da Silva, Matt Ratto, Myriam Delgado, Rodrigo Minetto, Daniel Silver, Thiago H Silva
arXiv AI
Sep 21

Do Personality-Tuned LLMs Make Better Social Agents?

The paper examines whether fine‑tuning large language models (LLMs) with personality‑labelled data improves their ability to act as socially interactive agents. Two small open‑weight LLMs were fine‑tuned on a corpus of personality‑labelled social media posts and dialogues, and the resulting models were evaluated in various social interaction scenarios by independent LLM judges. The findings show that the fine‑tuned models do not outperform their baseline counterparts in role‑playing personalities, though they offer comparable text quality and increased linguistic diversity for the Qwen models; low inter‑rater agreement limits confidence in the results, suggesting future work should focus on training data quality and domain alignment.

By Tim Krabbe, Xiaodan Shi
arXiv Computation and Language
Sep 3

PIVOTSBench: Evaluating Fine-Grained Interpersonal Relationship Reasoning in Multimodal Large Language Models

PIVOTSBench is a benchmark designed to assess multimodal large language models’ ability to reason about fine‑grained interpersonal relationships. It is constructed from Social‑IQ 2.0 and YouTube data and evaluates models on predicting bidirectional relationship dimensions grounded in psychology research. The benchmark also includes auxiliary tasks that test models’ capacity to identify and use critical visual cues, and it examines the impact of visual modalities, social role information, and different prediction settings on model performance.

By Shuxiang Zhang, Yiting Yin, Wenxuan Song, Yuhang Wu, Miao Liu