arXiv Computer Vision

HEDGE: A Calibrated Ensemble for A/H Recognition

Hugging Face Trending Papers
Jul 13

Simple Features and Honest Calibration for Ambivalence and Hesitancy Recognition in Video

We address ambivalence and hesitancy (A/H) recognition in the ABAW 2026 BAH Challenge: given a short interview video, predict whether the person shows signs of A/H. Our system combines affect-specialised text, audio, and visual representations with a small set of readable linguistic hesitation cues, fused by a reliability gate we call Affective Marker Fusion (AMF), and finished with a simple AP-weighted ensemble at a fixed decision threshold.

arXiv AI
Aug 24

Open-Weight Masked Introspection: Measuring What Language Models Can Report About Their Own Computation

The study investigates whether open‑weight language models can introspect on their own internal computations. Using the Open‑Weight Masked Introspection (OWMI) framework, researchers intervened on various internal components of eight models and asked them to report whether changes had occurred. Across 78,000 measurements, none of the models reliably distinguished real interventions from sham ones, with AUROC values essentially at chance. Why It Matters: The findings suggest that current open‑weight models lack the ability to audit their own internal states, highlighting a limitation for oversight that relies on a model’s self‑reporting.

By Emilio Ferrara
arXiv Computer Vision
Sep 2

Audio-Text Cross-Attention with Psycholinguistic Support Features for Ambivalence/Hesitancy Recognition

arXiv:2607.13345v2 Announce Type: replace Abstract: We present a frame-independent audio-text system for the 3rd Ambivalence/Hesitancy Video Recognition Challenge at the 11th Affective & Behavior Ana...

By Luiz F. B. F. Martins, Rodrigo W. Pisaia, Matheus M. Girardi, Isabella V. Berkembrock, Jo\~ao A. Almeida, Andre G. Hochuli, Rayson Laroca, Alceu S. Britto Jr
arXiv Machine Learning
Sep 3

Orthogonal Ensembles and Tested Explanations for Performer-Independent Body-Motion Emotion Recognition

The paper investigates 12‑class body‑only emotion recognition from skeleton motion using a leave‑performer‑out evaluation, where chance accuracy is 8.3% and a reproduced STGCN++ baseline scores 25.73% Macro‑F1. By ensembling eleven models with orthogonal error modes, the authors achieve 36.80% Macro‑F1, a 43% relative improvement over the baseline. They also introduce a tested explanation suite that demonstrates the ensemble’s decisions rely on motion‑grounded body‑region evidence, aligning strongly with Laban Movement Analysis attributes rather than classical kinematics, while showing diffuse temporal saliency.

By Naoto Nishida, Yoshio Ishiguro
arXiv Machine Learning
Sep 24

When Adaptation Hurts: Split Sensitivity and Person-Level Negative Transfer in Federated Wearable Onboarding

The paper evaluates six onboarding strategies for federated wearable models on five datasets using a leakage‑controlled protocol that fixes source checkpoints and separates calibration from evaluation. Results show that while average accuracy is high, person‑level performance can drop significantly, with some methods causing negative transfer for certain users. The study highlights that mean accuracy alone is insufficient and provides an auditable benchmark and failure map for future development.

By Rahil Aftab, Vineet Kumar Rakesh, Soumya Mazumdar, Tapas Samanta
arXiv AI
Jun 3

CoEval: Ranking Language Models for Custom Tasks Without Labeled Data or Trustworthy Benchmarks

arXiv:2606. 03650v1 Announce Type: cross Abstract: Choosing or ranking language models for a specific application is hardest when no task-specific labeled data exists, and standard public benchmarks cannot be trusted, their items having likely leaked into pretraining, so scores reflect memorization rather than fitness.

By Alexander Apartsin, Yehudit Aperstein
arXiv AI
Sep 17

Decodable but Misrouted: Sparse Features Uncover a Readout Gap in Vision-Language Models for Harmful Meme Detection

The paper investigates why large vision‑language models sometimes misclassify harmful memes, attributing failures to either missing internal evidence or poor routing of evidence to the output. Using sparse autoencoders, role‑conditioned probes, and causal interventions on Gemma‑3 and Qwen3.5, the authors show that sparse readouts consistently outperform native predictions across six harmful content benchmarks, revealing a readout gap that is largely due to routing rather than representation. The study also demonstrates that calibration‑only routing recovers most of the performance gap and that the issue persists across languages and is not solely driven by OCR signals.

By Girish A. Koushik, Diptesh Kanojia, Helen Treharne