arXiv AI By Shantanu Vispute, Aditya Mishra, Siddhartha Saxena

Who Said What, and Will It Be Remembered? Evaluating Persistent Speaker Attribution Across Meetings

Read the original on arXiv AI →

The Flow has not summarised this story yet — read it at arXiv AI.

arXiv AI
Sep 1

No Detectable Change in Side-Level WER from Prompt-Level Context: A Preregistered Ablation on a Production Oral-History Corpus

The study examined whether providing full prompt-level context to a large multimodal model would improve speech transcription accuracy on a production oral‑history corpus. Using a preregistered within‑item paired ablation, the authors found that adding context did not produce a detectable change in side‑level word error rate (WER) for either gpt‑4o‑transcribe or gemini‑2.5‑flash. The results suggest that context alone may not be sufficient to enhance aggregate transcription accuracy, and that finer‑grained, sequence‑aligned metrics are needed to evaluate such mechanisms.

By Theodore O. Cochran, Stephanie Dodson, Keith Nore
arXiv Computation and Language
Aug 24

The Voiceprint Fallacy: Why Voices Are Not Unique Biometric Imprints

The article critiques the notion that a person's voice is a stable, unique biometric trace—termed a voiceprint—by reviewing historical, forensic, and technological evidence. It argues that voices are highly dynamic and context-dependent, and that the voiceprint metaphor misrepresents probabilistic speaker information as a fixed identity marker. The authors emphasize that speaker recognition should account for within-speaker variability, domain mismatch, and synthetic manipulation rather than rely on an assumed stable voiceprint.

By Tianle Yang, Cuiling Zhang, Chengzhe Sun, Siwei Lyu, Phil Rose