arXiv AI

Who Said What, and Will It Be Remembered? Evaluating Persistent Speaker Attribution Across Meetings

arXiv AI
Sep 1

No Detectable Change in Side-Level WER from Prompt-Level Context: A Preregistered Ablation on a Production Oral-History Corpus

The study examined whether providing full prompt-level context to a large multimodal model would improve speech transcription accuracy on a production oral‑history corpus. Using a preregistered within‑item paired ablation, the authors found that adding context did not produce a detectable change in side‑level word error rate (WER) for either gpt‑4o‑transcribe or gemini‑2.5‑flash. The results suggest that context alone may not be sufficient to enhance aggregate transcription accuracy, and that finer‑grained, sequence‑aligned metrics are needed to evaluate such mechanisms.

By Theodore O. Cochran, Stephanie Dodson, Keith Nore
arXiv Computation and Language
Aug 24

The Voiceprint Fallacy: Why Voices Are Not Unique Biometric Imprints

The article critiques the notion that a person's voice is a stable, unique biometric trace—termed a voiceprint—by reviewing historical, forensic, and technological evidence. It argues that voices are highly dynamic and context-dependent, and that the voiceprint metaphor misrepresents probabilistic speaker information as a fixed identity marker. The authors emphasize that speaker recognition should account for within-speaker variability, domain mismatch, and synthetic manipulation rather than rely on an assumed stable voiceprint.

By Tianle Yang, Cuiling Zhang, Chengzhe Sun, Siwei Lyu, Phil Rose
arXiv Computation and Language
6d ago

Inference-Time Target Speaker Unlearning in LLM-Based Automatic Speech Recognition

The paper introduces a new target‑speaker unlearning task for automatic speech recognition (TSU‑ASR) that allows certain speakers to opt out of transcription while still indicating their presence. A lightweight Enrollment‑Conditioned Gating (ECG) module is added to a frozen dual‑stream speech LLM, enabling dynamic unlearning of new opt‑out speakers during inference. Experiments on AMI and AliMeeting datasets show significant drops in transcription accuracy for opt‑out speakers while preserving performance for retained speakers.

By Bo Su, Yueru Yan, Thai Le
arXiv Computation and Language
Sep 25

Interactive In-Meeting Speaker Correction with Human Feedback

The paper introduces an LLM-assisted system for correcting speaker attribution errors during meetings. It combines streaming ASR, diarization, and concise LLM-generated summaries to guide users in providing brief corrective feedback, which updates the transcript and adds online speaker enrollments. The approach includes mechanisms to accurately interpret user corrections and a simulation for large-scale evaluation, achieving significant reductions in DER and speaker substitution error on the AMI headset test set, with a pilot usability study highlighting further improvements.

By Xinlu He, Yiwen Guan, Badrivishal Paurana, Pitipat Kongsomjit, Zilin Dai, Jacob Whitehill
arXiv AI
Sep 24

Forget who you Forgot: Speaker Unlearning to Prevent Re-Identification in Zero-Shot Text-to-Speech

The paper introduces GUARD, a lightweight speaker identity unlearning framework designed to prevent re-identification in zero-shot text-to-speech systems. GUARD employs a learned speaker gate and speaker-agnostic activation steering on a frozen TTS backbone, optimizing steering vectors through group-relative reward optimization to reduce similarity to forgotten speakers while maintaining intelligibility and naturalness. Experiments on CosyVoice2 show that GUARD significantly lowers forget-speaker similarity and re-identification accuracy while preserving the ability to reproduce retained speakers.

By Hyoeun Kim, Yujun Lee, Kyuhong Shim
arXiv Machine Learning
Sep 4

VoxReason: Listener-Free Evaluation of Source-Grounded Speech Planning Before Synthesis

VoxReason introduces a listener‑free evaluation framework that measures whether a speech planning system’s delivery choices—such as pitch, energy, rate, pause, emphasis, and stance—are grounded in cited source records before any waveform is generated. The system outputs a source‑cited speaking plan and uses a deterministic verifier to check citation legality, slot agreement, unsupported states, schema validity, and counterfactual locality. Experiments on 1,440 source‑label cases show that simple slot accuracy can be misleading, while a 7B locality‑based repair model significantly improves plan‑slot accuracy and locality, and removing source records sharply reduces the grounded score. whyItMatters":"The framework provides a concrete, measurable way to ensure that expressive speech systems make source‑licensed planning decisions, addressing a source‑use failure that occurs before audio synthesis."

By Mengzhe Geng
arXiv AI
Sep 17

Beyond EER: Multi-Dimensional Evaluation of Information Leakage in Speaker De-Identification

The paper "Beyond EER: Multi-Dimensional Evaluation of Information Leakage in Speaker De-Identification" argues that evaluating speaker de-identification systems solely by Equal Error Rate (EER) is insufficient. It proposes a holistic framework using five complementary metrics—EER, soft biometric leakage score, cumulative match characteristic re-identification analysis, canonical correlation analysis with Procrustes embedding alignment, and intelligibility via word error rate and semantic similarity—to capture independent dimensions of information leakage. Experiments on five IARPA ARTS SDID systems show that these metrics reveal leakage that a single metric would miss.

By Seungmin Seo, Oleg Aulov, P. Jonathon Phillips, Kevin Mangold, Jonathan Eskin