Hugging Face Trending Papers

LLM-based Multimodal Personality Recognition via Facial Action Unit-Text Semantic Fusion

Personality recognition in asynchronous video interviews (AVIs) has become increasingly important due to their widespread adoption in modern recruitment. Existing approaches often rely on large language models (LLMs) to analyze textual responses of interviewees in AVI.

arXiv Computer Vision
Sep 21

Traits Run Deeper: Trait-Specific Asymmetric Fusion for Multimodal Personality Assessment

Traits Run Deeper introduces a personality assessment framework that tailors multimodal fusion to each trait dimension. It comprises a Multimodal Foundation Representation module that uses psychology-informed semantic templates, a Trait-Specific Modality Fusion module that asymmetrically fuses modalities to reduce cross‑modal interference, and a Distribution‑Calibrated Personality Regression module that corrects label imbalance. The approach achieves a ~25% reduction in mean squared error on the AVI Challenge 2026 validation set and wins the Personality Assessment Track.

By Jia Li, Qian Chen, Wei Wang, Xinyu Li, Zhenzhen Hu, Dongsheng Shao, Richang Hong, Meng Wang
arXiv AI
Jun 11

Frozen Multimodal Embeddings for Personality and Cognitive Ability Assessment in Asynchronous Video Interviews

arXiv:2606. 11930v1 Announce Type: cross Abstract: Predicting psychological traits from asynchronous video interviews (AVIs) is a challenging multimodal learning problem because labeled datasets are limited while each response contains high-dimensional visual, acoustic, and verbal signals.

By Kuo-En Hung, Hung-Yue Suen, Shih-Ching Yeh, Hsiang-Wen Wang
arXiv AI
Jun 12

Frozen Multimodal Embeddings for AI-Assisted Interview Assessment of Personality and Cognitive Ability

arXiv:2606. 11930v2 Announce Type: replace-cross Abstract: Predicting psychological traits from asynchronous video interviews (AVIs) is a challenging problem in AI-assisted interview assessment because labeled datasets are limited while each response contains high-dimensional visual, acoustic, and verbal signals.

By Kuo-En Hung, Hung-Yue Suen, Shih-Ching Yeh, Hsiang-Wen Wang
arXiv AI
Jun 10

Modeling Complex Behaviors: Multi-Personality Composition and Dynamic Switching in Vision-Language Models

arXiv:2606. 11074v1 Announce Type: cross Abstract: With the widespread deployment of Multimodal Large Language Models (MLLMs) in social interaction, understanding and controlling their behavior under complex personality conditions is essential.

By Peiqi Jia (Xi'an Jiaotong University), Haonan Jia (Beihang University), Ziqi Miao (Beihang University), Linkang Du (Xi'an Jiaotong University), Yuntao Wang (Xi'an Jiaotong University), Zhou Su (Xi'an Jiaotong University)
arXiv AI
Aug 19

Beyond BFI: The CSI for Enhanced Reliability and Validity in Evaluating LLM Personality Traits

The paper introduces the Core Sentiment Inventory (CSI), a new personality trait evaluation tool for large language models (LLMs) that addresses reliability and validity issues found in existing methods like the Big Five Inventory (BFI). CSI is designed specifically for LLMs, supports both English and Chinese, and provides detailed psychological portraits of model behavior. Experiments show that CSI captures nuanced behavioral patterns, improves reliability, and correlates strongly (above 0.85) with real-world LLM outputs.

By Huanhuan Ma, Haisong Gong, Xiaoyuan Yi, Xing Xie, Philip S. Yu, Dongkuan Xu
arXiv Computer Vision
Aug 28

HUG-VIS: A Multimodal Benchmark for Human-centered Understanding and Generation in Visual Intelligence

HUG‑VIS is a unified multimodal benchmark for human‑centered visual intelligence, comprising 8,400 half‑body videos of 30 professional actors performing 280 emotion‑action prompts in Mandarin. The dataset provides synchronized video, audio, text, and alpha mattes for four tasks—human emotion recognition, video generation, voice cloning, and video matting—allowing evaluation of both open‑ and closed‑source models under a zero‑shot protocol. Results reveal that linguistic cues dominate emotion recognition, visual affect is weakest, and that automatic metrics and human judgments diverge in generation and cloning tasks, while motion‑related boundary fidelity remains a key challenge for matting.

By Fei Ma, Zebang Cheng, Minghui Li, Hongbo Xu, Yuyong Tan, Yihua Shao, Hanling Wang, Zhou Liu, Yuqing Gao, Dong Wang, Long Ma, Laizhong Cui, Nicu Sebe, Qi Tian