arXiv Computer Vision By Fei Ma, Zebang Cheng, Minghui Li, Hongbo Xu, Yuyong Tan, Yihua Shao, Hanling Wang, Zhou Liu, Yuqing Gao, Dong Wang, Long Ma, Laizhong Cui, Nicu Sebe, Qi Tian

HUG-VIS: A Multimodal Benchmark for Human-centered Understanding and Generation in Visual Intelligence

Read the original on arXiv Computer Vision →

HUG‑VIS is a unified multimodal benchmark for human‑centered visual intelligence, comprising 8,400 half‑body videos of 30 professional actors performing 280 emotion‑action prompts in Mandarin. The dataset provides synchronized video, audio, text, and alpha mattes for four tasks—human emotion recognition, video generation, voice cloning, and video matting—allowing evaluation of both open‑ and closed‑source models under a zero‑shot protocol. Results reveal that linguistic cues dominate emotion recognition, visual affect is weakest, and that automatic metrics and human judgments diverge in generation and cloning tasks, while motion‑related boundary fidelity remains a key challenge for matting.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Computer Vision.

arXiv AI
Jun 11

Frozen Multimodal Embeddings for Personality and Cognitive Ability Assessment in Asynchronous Video Interviews

arXiv:2606. 11930v1 Announce Type: cross Abstract: Predicting psychological traits from asynchronous video interviews (AVIs) is a challenging multimodal learning problem because labeled datasets are limited while each response contains high-dimensional visual, acoustic, and verbal signals.

By Kuo-En Hung, Hung-Yue Suen, Shih-Ching Yeh, Hsiang-Wen Wang
arXiv AI
Jun 12

Frozen Multimodal Embeddings for AI-Assisted Interview Assessment of Personality and Cognitive Ability

arXiv:2606. 11930v2 Announce Type: replace-cross Abstract: Predicting psychological traits from asynchronous video interviews (AVIs) is a challenging problem in AI-assisted interview assessment because labeled datasets are limited while each response contains high-dimensional visual, acoustic, and verbal signals.

By Kuo-En Hung, Hung-Yue Suen, Shih-Ching Yeh, Hsiang-Wen Wang
Hugging Face Trending Papers
Jul 23

MVEI & EmObserver: Empowering MLLM-Oriented Visual Emotional Intelligence via Emotion Statement Judgement

Affective Image Content Analysis (AICA) aims to recognize and understand emotions elicited by visual content, representing an indispensable step toward Artificial General Intelligence (AGI). However, despite the rapid progress of Multimodal Large Language Models (MLLMs), systematic evaluation of their visual emotional intelligence remains largely absent from recent model releases.

Hugging Face Trending Papers
Aug 11

E$^3$mo-Bench: A Scalable Benchmark for Multimodal Evoked and Expressed Emotion Understanding via Bayesian Pairwise Alignment

Understanding both expressed and evoked emotions is critical for multimodal large language models (MLLMs) to achieve comprehensive affect-aware interactions. However, existing benchmarks typically examine expressed and evoked emotions in isolation or are constrained to coarse-grained and incomplete affective characterizations.