arXiv Machine Learning

Orthogonal Ensembles and Tested Explanations for Performer-Independent Body-Motion Emotion Recognition

The paper investigates 12‑class body‑only emotion recognition from skeleton motion using a leave‑performer‑out evaluation, where chance accuracy is 8.3% and a reproduced STGCN++ baseline scores 25.73% Macro‑F1. By ensembling eleven models with orthogonal error modes, the authors achieve 36.80% Macro‑F1, a 43% relative improvement over the baseline. They also introduce a tested explanation suite that demonstrates the ensemble’s decisions rely on motion‑grounded body‑region evidence, aligning strongly with Laban Movement Analysis attributes rather than classical kinematics, while showing diffuse temporal saliency.

arXiv Computer Vision
Aug 27

HEDGE: A Calibrated Ensemble for A/H Recognition

arXiv:2607.12176v2 Announce Type: replace Abstract: Ambivalence and hesitancy (A/H) undermine digital behaviour-change interventions, and recognizing them automatically from video is the goal of the...

By Josep Cabacas-Maso, Ismael Benito-Altamirano, Carles Ventura
arXiv Machine Learning
Aug 31

What Do Interaction Representations Actually Measure? Pre-Event Separability in Weakly-Supervised Violence Detection

The paper investigates whether detailed articulated human pose provides more discriminative power than coarse spatial relationships for early violence detection. By fixing the downstream pipeline and comparing five interaction representations—including bounding‑box geometry, handcrafted pose analogues, enriched pose descriptors, and a learned joint encoder—the study finds that pose‑based representations do not outperform coarse geometry. When visual encoders are frozen and evaluated on larger datasets, person‑crop appearance and whole‑frame context outperform geometry, but cropping to interacting people offers no advantage over encoding the entire frame. The authors further demonstrate that pre‑onset frames contain source‑related artifacts (e.g., title cards, watermarks) that contribute significantly to discrimination, suggesting that benchmark performance may reflect these artifacts rather than true event evidence.

By Parishruthi Ganesh
arXiv Machine Learning
Jun 16

Faithful Action-unit Causal Reasoning for Counterfactually Faithful Emotion Explanations

arXiv:2606. 15779v1 Announce Type: cross Abstract: Multimodal models can name the action units (AUs) behind a facial emotion, but their AU->emotion rationales are typically plausible rather than faithful: nothing forces the AUs a model invokes to be the AUs that actually drive its prediction.

By Van Thong Huynh, Hong Hai Nguyen, Thuy Pham, Trong Nghia Nguyen, Soo-Hyung Kim
arXiv Machine Learning
Sep 16

EMODY Flow: Emotion-Aware Audio-Driven Full-Body Motion Generation

EMODY Flow is a lightweight flow‑matching framework that generates synchronized full‑body motion and facial expressions conditioned on speech and emotion. It attaches to a frozen Qwen‑3 Omni model, reusing its audio codecs to drive two parallel DiT generators for SMPL‑X body pose and FLAME facial expressions. An auxiliary emotion classifier at training time restores emotion sensitivity, enabling EMODY Flow to achieve state‑of‑the‑art gesture quality on BEAT2 and zero‑shot facial animation on TFHP, with significant improvements in FGD, Beat Correlation, and Diversity metrics.

By Harsh Kumar Agarwal, Xavier Alameda-Pineda, Olivier Perrotin
Hugging Face Trending Papers
Jul 30

A Montage-Agnostic Encoder for Calibration-Light Cross-User Gesture Recognition from Surface Electromyography

Pattern-recognition control promises a myoelectric prosthesis that responds to many intended gestures rather than one or two, but the promise has stayed in the laboratory. A recogniser trained on one person rarely transfers to the next, and useful performance usually demands a fresh round of labelled calibration from the end user.

arXiv AI
Sep 4

Cross-Dataset Transfer and Reliability of Explainable Artificial Intelligence for RhythmFormer Remote Photoplethysmography

The study evaluates explainable AI methods for Remote Photoplethysmography (rPPG) using the RhythmFormer model across two datasets (NCKU-rPPG and UBFC-rPPG). Quantitative metrics—skin coverage and the Salience-guided Faithfulness Coefficient (SaCo)—were applied to four explanation techniques (raw attention, rollout, attention flow, and Beyond Intuition). Beyond Intuition consistently achieved the highest coverage and SaCo, while other methods showed weak or opposite correlations with performance measures, and performance degraded notably under low illumination (40 lux).

By Louis Chen, Torbj\"orn E. M. Nordling
arXiv AI
Aug 24

Aggregate, Don't Adapt: Subject-Level Posterior Aggregation and Transductive Calibration for Cross-Site Parkinsonian Gait Severity

The paper reports the winning solution to the MoCha 2026 Parkinsonian Gait Benchmark, achieving a macro‑F1 score of 0.6945 on unseen clinical sites. The approach relies on a frozen public motion encoder followed by a single 4×512 linear layer, and gains are largely attributed to three key steps: exact replication of the benchmark’s head recipe, averaging per‑walk posteriors at the subject level, and a label‑free transductive calibration of feature means and decision thresholds. Extensive ablation studies show that fine‑tuning the encoder or using alternative encoders does not improve performance, and the subject‑level aggregation is identified as the primary contributor to the top score.

By Junlong Shen