One-Frame Calibration with Siamese Network in Facial Action Unit Recognition
arXiv:2409. 00240v2 Announce Type: replace-cross Abstract: Automatic facial action unit (AU) recognition is used widely in facial expression analysis.
The paper introduces a disagreement‑aware dynamic facial expression recognition framework that directly learns from raw annotator vote vectors using a Dirichlet‑Multinomial likelihood, preserving both predictive mean and scale‑sensitive supervision. It adds an ambiguity head to estimate annotation entropy for unseen clips and employs a Chow‑style reject rule that integrates ambiguity, vacuity, temporal instability, and input quality for selective prediction. On the DFEW benchmark, the method maintains recognition accuracy while cutting expected calibration error by 30 % and area‑under‑risk‑curve by 15 %, with predicted ambiguity correlating 0.52 (Spearman) with true annotation entropy, and these gains transfer to FERV39k and hold under identity‑ and movie‑disjoint splits.
arXiv:2409. 00240v2 Announce Type: replace-cross Abstract: Automatic facial action unit (AU) recognition is used widely in facial expression analysis.
arXiv:2607. 12774v1 Announce Type: cross Abstract: This article presents our results for the 11th Affective Behavior Analysis in-the-Wild (ABAW) competition.
arXiv:2604. 15280v2 Announce Type: replace-cross Abstract: Understanding emotions is a fundamental ability for intelligent systems to be able to interact with humans.
arXiv:2606. 27536v1 Announce Type: cross Abstract: Speech emotion recognition (SER) often relies on hard consensus labels that collapse annotator disagreement.
Chehre is an emoji‑prompted video dataset designed to study perceptual flexibility in video language models. It contains 2,111 videos of 203 participants expressing 40 facial emojis, with each video annotated by about 30 perceivers, yielding 1,242 annotators in total. The dataset introduces a new task—distributional expression recognition—that evaluates a model’s ability to reproduce the variation seen in human annotations, and shows that persona prompting can shift model perception to better match human variability.
Conventional face recognition relies on static appearance cues and degrades in unconstrained settings with expression variation, occlusion, and poor lighting. We hypothesize that audiovisual expression dynamics carry identity-discriminative information complementary to static appearance, and that extracting this signal requires multimodal representations robust to the variable input quality of in-the-wild video.
arXiv:2606. 02679v1 Announce Type: new Abstract: Multimodal systems often benefit from combining information across language, sound, and visual streams, but this benefit is not guaranteed.
The paper tackles the challenge of predicting student engagement from online tutoring videos, noting that engagement is a complex, multidimensional construct influenced by behavioral, emotional, and cognitive states. By analyzing the CASED dataset, the authors highlight the difficulty posed by high inter‑person variability and subjective annotations. They propose a multimodal framework that fuses implicit spatiotemporal features from pretrained video, audio, and image encoders with structured behavioral cues such as head pose, gaze, facial action units, emotion, and wavelet‑based audio features, integrating them via a Perceiver IO bottleneck and modeling participant personalities with variational posteriors. The system employs evidential regression and spectral‑normalized Gaussian process classification heads to provide uncertainty‑aware predictions, achieving competitive performance on the CASED challenge test set while offering well‑calibrated uncertainty metrics.
The paper presents Test-Time Adaptation via Cache Personalization (TTA‑CaP), a gradient‑free, cache‑based method that personalizes vision‑language models for facial expression recognition in videos. TTA‑CaP uses three complementary caches—a personalized static cache, a positive target cache, and a negative target cache—controlled by a tri‑gate mechanism to prevent corruption and provide robust subject‑matched evidence. Experiments on BioVid, StressID, and BAH datasets show that TTA‑CaP outperforms state‑of‑the‑art test‑time adaptation methods while keeping computational and memory overhead low.
arXiv:2609.15608v1 Announce Type: cross Abstract: Detecting sexism on the internet is a fundamentally subjective task; our team, VANGUARD, addresses this challenge in the EXIST 2026 Task 2 by proposi...
arXiv:2607. 21820v1 Announce Type: cross Abstract: Audio deepfake detectors are trained to distinguish genuine speech from synthetic speech and often perform well on standard benchmarks.
arXiv:2511. 14117v2 Announce Type: replace Abstract: Supervised classifiers output a distribution over classes but are typically trained against a single label obtained by collapsing multiple annotators into a majority vote.