arXiv:2609.37297v1 Announce Type: new
Abstract: Cross-skeleton motion generation trains generative models to carry action structure and motion intention from one body to another. Yet a target motion...
By Zhiyuan Li, Wenyan Yang, Pekka Marttinen, Joni Pajarinen
arXiv:2607.12176v2 Announce Type: replace
Abstract: Ambivalence and hesitancy (A/H) undermine digital behaviour-change interventions, and recognizing them automatically from video is the goal of the...
By Josep Cabacas-Maso, Ismael Benito-Altamirano, Carles Ventura
The paper investigates whether detailed articulated human pose provides more discriminative power than coarse spatial relationships for early violence detection. By fixing the downstream pipeline and comparing five interaction representations—including bounding‑box geometry, handcrafted pose analogues, enriched pose descriptors, and a learned joint encoder—the study finds that pose‑based representations do not outperform coarse geometry. When visual encoders are frozen and evaluated on larger datasets, person‑crop appearance and whole‑frame context outperform geometry, but cropping to interacting people offers no advantage over encoding the entire frame. The authors further demonstrate that pre‑onset frames contain source‑related artifacts (e.g., title cards, watermarks) that contribute significantly to discrimination, suggesting that benchmark performance may reflect these artifacts rather than true event evidence.
By Parishruthi Ganesh
arXiv:2606. 15779v1 Announce Type: cross Abstract: Multimodal models can name the action units (AUs) behind a facial emotion, but their AU->emotion rationales are typically plausible rather than faithful: nothing forces the AUs a model invokes to be the AUs that actually drive its prediction.
By Van Thong Huynh, Hong Hai Nguyen, Thuy Pham, Trong Nghia Nguyen, Soo-Hyung Kim
EMODY Flow is a lightweight flow‑matching framework that generates synchronized full‑body motion and facial expressions conditioned on speech and emotion. It attaches to a frozen Qwen‑3 Omni model, reusing its audio codecs to drive two parallel DiT generators for SMPL‑X body pose and FLAME facial expressions. An auxiliary emotion classifier at training time restores emotion sensitivity, enabling EMODY Flow to achieve state‑of‑the‑art gesture quality on BEAT2 and zero‑shot facial animation on TFHP, with significant improvements in FGD, Beat Correlation, and Diversity metrics.
By Harsh Kumar Agarwal, Xavier Alameda-Pineda, Olivier Perrotin
Pattern-recognition control promises a myoelectric prosthesis that responds to many intended gestures rather than one or two, but the promise has stayed in the laboratory. A recogniser trained on one person rarely transfers to the next, and useful performance usually demands a fresh round of labelled calibration from the end user.
arXiv:2607. 27568v1 Announce Type: new Abstract: Recognition accuracy obtained during a recording session does not persist when a user puts on the electrodes again after the electrodes had previously been removed.
By Jethro Odeyemi, W. J. Zhang
The study evaluates explainable AI methods for Remote Photoplethysmography (rPPG) using the RhythmFormer model across two datasets (NCKU-rPPG and UBFC-rPPG). Quantitative metrics—skin coverage and the Salience-guided Faithfulness Coefficient (SaCo)—were applied to four explanation techniques (raw attention, rollout, attention flow, and Beyond Intuition). Beyond Intuition consistently achieved the highest coverage and SaCo, while other methods showed weak or opposite correlations with performance measures, and performance degraded notably under low illumination (40 lux).
By Louis Chen, Torbj\"orn E. M. Nordling
arXiv:2607. 20820v1 Announce Type: new Abstract: Body-based emotion recognition is important for real-time affective systems, but graph-based skeleton models can be computationally expensive.
By Christian Arzate Cruz, Stefanos Gkikas, Houshyar Asadi
arXiv:2609.13308v1 Announce Type: cross
Abstract: A companion evaluation found that naming the target part in a manipulation prompt increased action accuracy by 0.32-0.63 across eight vision-language...
By Sarthak Sattigeri
The paper reports the winning solution to the MoCha 2026 Parkinsonian Gait Benchmark, achieving a macro‑F1 score of 0.6945 on unseen clinical sites. The approach relies on a frozen public motion encoder followed by a single 4×512 linear layer, and gains are largely attributed to three key steps: exact replication of the benchmark’s head recipe, averaging per‑walk posteriors at the subject level, and a label‑free transductive calibration of feature means and decision thresholds. Extensive ablation studies show that fine‑tuning the encoder or using alternative encoders does not improve performance, and the subject‑level aggregation is identified as the primary contributor to the top score.
By Junlong Shen
arXiv:2607. 27565v1 Announce Type: new Abstract: Pattern-recognition control promises a myoelectric prosthesis that responds to many intended gestures rather than one or two, but the promise has stayed in the laboratory.
By Jethro Odeyemi, W. J. Zhang