Who Said What, and Will It Be Remembered? Evaluating Persistent Speaker Attribution Across Meetings
Read the original on arXiv AI →The Flow has not summarised this story yet — read it at arXiv AI.
The Flow has not summarised this story yet — read it at arXiv AI.
arXiv:2608.22196v1 Announce Type: cross Abstract: While cascaded multi-talker ASR (MT-ASR) leverages state-of-the-art foundation models, its performance is often capped by speaker leakage during sepa...
arXiv:2607. 21820v1 Announce Type: cross Abstract: Audio deepfake detectors are trained to distinguish genuine speech from synthetic speech and often perform well on standard benchmarks.
The study examined whether providing full prompt-level context to a large multimodal model would improve speech transcription accuracy on a production oral‑history corpus. Using a preregistered within‑item paired ablation, the authors found that adding context did not produce a detectable change in side‑level word error rate (WER) for either gpt‑4o‑transcribe or gemini‑2.5‑flash. The results suggest that context alone may not be sufficient to enhance aggregate transcription accuracy, and that finer‑grained, sequence‑aligned metrics are needed to evaluate such mechanisms.
Full-duplex speech models require training data that preserves turn-taking, overlap, interruption, and backchannel behavior, yet these signals are entangled across speakers in noisy real-world recordi...
arXiv:2608.28916v1 Announce Type: new Abstract: Automatic speech recognition (ASR) systems are commonly evaluated with word error rate (WER), yet many voice workflows depend on exact written values f...
The article critiques the notion that a person's voice is a stable, unique biometric trace—termed a voiceprint—by reviewing historical, forensic, and technological evidence. It argues that voices are highly dynamic and context-dependent, and that the voiceprint metaphor misrepresents probabilistic speaker information as a fixed identity marker. The authors emphasize that speaker recognition should account for within-speaker variability, domain mismatch, and synthetic manipulation rather than rely on an assumed stable voiceprint.