arXiv AI By Mohammed Hafsati, Ahmed Loughzali

Multi-Party Backchannel Prediction: a Diagnosis, a Benchmark, and a Ceiling

Read the original on arXiv AI →

The paper introduces a multi‑party backchannel prediction benchmark built from the AMI meeting corpus, featuring 682 masked‑listener views, 190 speakers, and 18,697 backchannel events. A state‑of‑the‑art dyadic model performs at chance when applied zero‑shot to meetings, but its frozen acoustic features are still informative, and retraining improves performance to an AUROC of 0.751. The study reveals that listener conditioning helps only for listeners seen during training, that speaker identity is entangled with useful cues, and that backchannel rates vary significantly across individuals, prompting the authors to report both AUROC and event‑F1 metrics. whyItMatters":"The benchmark and evaluation tools provide a standardized, person‑disjoint testbed for advancing multi‑party backchannel prediction research."

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Computation and Language
6d ago

Asymmetric Classifier-Free Guidance for Target-Speaker ASR

The paper introduces asymmetric classifier‑free guidance (CFG) for target‑speaker ASR using Whisper, where a speaker‑conditioned branch predicts the target transcript and a speaker‑unconditioned branch predicts serialized multi‑speaker transcripts. CFG modulates the influence of speaker conditioning during decoding via a single guidance scale, which is first set globally on development data and then refined per utterance by a lightweight encoder‑based predictor while keeping the recognition model fixed. The resulting system yields up to 21.8% relative WER reduction over a condition‑only baseline and 5.6% over standard conditional decoding under domain shifts.

By Yiwen Guan, Jacob Whitehill
arXiv AI
Aug 24

Do SpeechLMs Hear Their Own Opinions? Diagnosing and Mitigating Previous-Belief Contamination in Streaming Emotion Understanding

The paper investigates how streaming emotion recognition models can be misled by their own prior predictions, a problem termed previous-belief contamination (PBC). Using a counterfactual diagnostic on CREMA-D-Stream, the authors show that feeding a model’s previous emotion label into its current prediction can drastically lower accuracy and flip many predictions, with the effect varying by label. To mitigate PBC, they propose EmoUpdate, a training‑free framework that isolates current audio perception from historical context through a prior‑blind firewall, a causal belief filter, and a decontamination operator, achieving significant gains across multiple SpeechLMs and benchmarks.

By Haoyue Liu, Zhichao Wang, Ye Chen, Haonan Deng, Xiaoying Tang