We address ambivalence and hesitancy (A/H) recognition in the ABAW 2026 BAH Challenge: given a short interview video, predict whether the person shows signs of A/H. Our system combines affect-specialised text, audio, and visual representations with a small set of readable linguistic hesitation cues, fused by a reliability gate we call Affective Marker Fusion (AMF), and finished with a simple AP-weighted ensemble at a fixed decision threshold.
arXiv:2607. 17384v1 Announce Type: new Abstract: This paper provides an experimentally verified formal law for calculating the uplift that diversity of thought provides in Large Language Model (LLM) ensembles.
By Junade Ali
The study investigates whether open‑weight language models can introspect on their own internal computations. Using the Open‑Weight Masked Introspection (OWMI) framework, researchers intervened on various internal components of eight models and asked them to report whether changes had occurred. Across 78,000 measurements, none of the models reliably distinguished real interventions from sham ones, with AUROC values essentially at chance.
Why It Matters: The findings suggest that current open‑weight models lack the ability to audit their own internal states, highlighting a limitation for oversight that relies on a model’s self‑reporting.
By Emilio Ferrara
arXiv:2607. 12774v1 Announce Type: cross Abstract: This article presents our results for the 11th Affective Behavior Analysis in-the-Wild (ABAW) competition.
By Aleksei Bakin, Andrey V. Savchenko
arXiv:2608.30456v1 Announce Type: new
Abstract: We compare six self-supervised pretext tasks for infant cry analysis under a fixed budget, meaning the same compact encoder of 1.17M parameters, the sa...
By Luigi Simeone
arXiv:2609.00654v1 Announce Type: new
Abstract: We describe the SciTrue team's participation in both subtasks of the NTCIR-19 SciClaimEval task~\cite{sciclaimeval}, which asks systems to verify scien...
By Qiming Bao, Ne\c{s}et \"Ozkan Tan, Siyuan Wang, Mark Gahegan
arXiv:2607. 25961v1 Announce Type: cross Abstract: Ambivalence and hesitancy (A/H) are conflicting affective states that precede the delay or abandonment of health behaviour change.
By Podakanti Satyajith Chary, Barath Parthiban, Pranesh Velmurugan, Adeeba Khan, Nagarajan Ganapathy
arXiv:2607.13345v2 Announce Type: replace
Abstract: We present a frame-independent audio-text system for the 3rd Ambivalence/Hesitancy Video Recognition Challenge at the 11th Affective & Behavior Ana...
By Luiz F. B. F. Martins, Rodrigo W. Pisaia, Matheus M. Girardi, Isabella V. Berkembrock, Jo\~ao A. Almeida, Andre G. Hochuli, Rayson Laroca, Alceu S. Britto Jr
The paper investigates 12‑class body‑only emotion recognition from skeleton motion using a leave‑performer‑out evaluation, where chance accuracy is 8.3% and a reproduced STGCN++ baseline scores 25.73% Macro‑F1. By ensembling eleven models with orthogonal error modes, the authors achieve 36.80% Macro‑F1, a 43% relative improvement over the baseline. They also introduce a tested explanation suite that demonstrates the ensemble’s decisions rely on motion‑grounded body‑region evidence, aligning strongly with Laban Movement Analysis attributes rather than classical kinematics, while showing diffuse temporal saliency.
By Naoto Nishida, Yoshio Ishiguro
The paper evaluates six onboarding strategies for federated wearable models on five datasets using a leakage‑controlled protocol that fixes source checkpoints and separates calibration from evaluation. Results show that while average accuracy is high, person‑level performance can drop significantly, with some methods causing negative transfer for certain users. The study highlights that mean accuracy alone is insufficient and provides an auditable benchmark and failure map for future development.
By Rahil Aftab, Vineet Kumar Rakesh, Soumya Mazumdar, Tapas Samanta
arXiv:2606. 03650v1 Announce Type: cross Abstract: Choosing or ranking language models for a specific application is hardest when no task-specific labeled data exists, and standard public benchmarks cannot be trusted, their items having likely leaked into pretraining, so scores reflect memorization rather than fitness.
By Alexander Apartsin, Yehudit Aperstein
The paper investigates why large vision‑language models sometimes misclassify harmful memes, attributing failures to either missing internal evidence or poor routing of evidence to the output. Using sparse autoencoders, role‑conditioned probes, and causal interventions on Gemma‑3 and Qwen3.5, the authors show that sparse readouts consistently outperform native predictions across six harmful content benchmarks, revealing a readout gap that is largely due to routing rather than representation. The study also demonstrates that calibration‑only routing recovers most of the performance gap and that the issue persists across languages and is not solely driven by OCR signals.
By Girish A. Koushik, Diptesh Kanojia, Helen Treharne