HUG‑VIS is a unified multimodal benchmark for human‑centered visual intelligence, comprising 8,400 half‑body videos of 30 professional actors performing 280 emotion‑action prompts in Mandarin. The dataset provides synchronized video, audio, text, and alpha mattes for four tasks—human emotion recognition, video generation, voice cloning, and video matting—allowing evaluation of both open‑ and closed‑source models under a zero‑shot protocol. Results reveal that linguistic cues dominate emotion recognition, visual affect is weakest, and that automatic metrics and human judgments diverge in generation and cloning tasks, while motion‑related boundary fidelity remains a key challenge for matting.
By Fei Ma, Zebang Cheng, Minghui Li, Hongbo Xu, Yuyong Tan, Yihua Shao, Hanling Wang, Zhou Liu, Yuqing Gao, Dong Wang, Long Ma, Laizhong Cui, Nicu Sebe, Qi Tian
arXiv:2608.22993v1 Announce Type: new
Abstract: Students increasingly use LLMs as tutors for coursework and problem solving. Little is known about the level of assistance LLMs provide when students u...
By Suhyeon Lee, Juneha Baek, Jaehyeong Park, Donghyuk Shin
The paper tackles the challenge of predicting student engagement from online tutoring videos, noting that engagement is a complex, multidimensional construct influenced by behavioral, emotional, and cognitive states. By analyzing the CASED dataset, the authors highlight the difficulty posed by high inter‑person variability and subjective annotations. They propose a multimodal framework that fuses implicit spatiotemporal features from pretrained video, audio, and image encoders with structured behavioral cues such as head pose, gaze, facial action units, emotion, and wavelet‑based audio features, integrating them via a Perceiver IO bottleneck and modeling participant personalities with variational posteriors. The system employs evidential regression and spectral‑normalized Gaussian process classification heads to provide uncertainty‑aware predictions, achieving competitive performance on the CASED challenge test set while offering well‑calibrated uncertainty metrics.
By Alperen Kantarci, Visvanathan Ramesh, Gemma Roig
Automated classroom engagement recognition holds substantial promise for scalable learning analytics, yet the suitability of modern Vision-Language Models (VLMs) for this task under zero-shot conditions remains largely unexplored. We present a systematic benchmark that evaluates five widely-used VLMs: CLIP, BLIP-VQA, GPT-4o, LLaVA-1.
The paper introduces Disengagement-Aware Student Simulators (DAS2), a protocol that models five learner-engagement states—engaged, gaming, wheel-spinning, off-task, and mixed—to evaluate AI tutor performance before deployment. Using annotated tutoring sessions from ASSISTments09, DAS2’s rule-based labels matched human consensus in 81% of cases, and conditioning simulations on intended states narrowed the correctness-rate gap between simulated and authentic sessions for gaming and wheel-spinning behaviors. The study also compares five AI tutors across these states, finding stable relative rankings but state-specific performance differences, and notes that automated evaluation does not fully align with human judgment.
By Xianghui Meng, Jionghao Lin
Personality recognition in asynchronous video interviews (AVIs) has become increasingly important due to their widespread adoption in modern recruitment. Existing approaches often rely on large language models (LLMs) to analyze textual responses of interviewees in AVI.
arXiv:2606. 16428v1 Announce Type: cross Abstract: Effective personalized AI-assisted learning demands systems that can not only generate accurate learner-specific educational materials, but also dynamically adapt their instruction to diverse learners.
By Jaward Sesay, Yue Yu, Siwei Dong, Yemin Shi, Guangyao Chen, B\"orje F. Karlsson
arXiv:2609.22778v1 Announce Type: new
Abstract: Multimodal large language models (MLLMs) are increasingly used as evaluators, yet their reliability in professional assessment tasks that require exper...
By Yuhan Lu, Yi Yao, Hua Shen, Katie Aafjes-van Doorn, Zhaonan Wang
arXiv:2606. 01584v1 Announce Type: cross Abstract: Conversational tutoring agents have been shown to improve learning engagement and student outcomes, and large language models (LLMs) are increasingly used in these systems to provide scalable, personalized feedback.
By Aitor Arronte Alvarez, Naiyi Xie Fincham
arXiv:2608. 10448v1 Announce Type: new Abstract: Multimodal emotion recognition in conversation (MERC) requires understanding complex interactions between verbal and non-verbal cues.
By Sujung Oh, Jung Uk Kim, Sangmin Lee
The paper presents a tutoring platform that combines a generative AI chatbot with a reinforcement learning algorithm to adaptively sequence practice problems for students learning Python. In a five‑month field study across ten high schools, the adaptive sequencing improved unassisted final exam performance by 0.15 standard deviations, with mediation analysis indicating that higher engagement drove the gains. The study demonstrates that signals from student‑chatbot interactions can be leveraged to personalize and optimize learning at scale.
By Angel Tsai-Hsuan Chung, Botong Zhang, Ling-Chieh Kung, Hamsa Bastani, Osbert Bastani
arXiv:2608. 03952v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used to provide conversational practice for English-as-a-second-language (ESL) learners.
By Dongjie Yang, Siyan Lin, Leixian Shen, Rui Sheng, Huamin Qu, Zixin Chen