arXiv Computation and Language

The Voiceprint Fallacy: Why Voices Are Not Unique Biometric Imprints

The article critiques the notion that a person's voice is a stable, unique biometric trace—termed a voiceprint—by reviewing historical, forensic, and technological evidence. It argues that voices are highly dynamic and context-dependent, and that the voiceprint metaphor misrepresents probabilistic speaker information as a fixed identity marker. The authors emphasize that speaker recognition should account for within-speaker variability, domain mismatch, and synthetic manipulation rather than rely on an assumed stable voiceprint.

arXiv AI
2d ago

Beyond EER: Multi-Dimensional Evaluation of Information Leakage in Speaker De-Identification

The paper "Beyond EER: Multi-Dimensional Evaluation of Information Leakage in Speaker De-Identification" argues that evaluating speaker de-identification systems solely by Equal Error Rate (EER) is insufficient. It proposes a holistic framework using five complementary metrics—EER, soft biometric leakage score, cumulative match characteristic re-identification analysis, canonical correlation analysis with Procrustes embedding alignment, and intelligibility via word error rate and semantic similarity—to capture independent dimensions of information leakage. Experiments on five IARPA ARTS SDID systems show that these metrics reveal leakage that a single metric would miss.

By Seungmin Seo, Oleg Aulov, P. Jonathon Phillips, Kevin Mangold, Jonathan Eskin
arXiv Computer Vision
Sep 11

Leveraging Avatar Fingerprinting: A Multi-Generator Photorealistic Talking-Head Public Database and Benchmark

The paper introduces AVAPrintDB, a new public multi‑generator talking‑head avatar database designed for avatar fingerprinting, comprising data from two audiovisual corpora and three state‑of‑the‑art generators (GAGAvatar, LivePortrait, HunyuanPortrait). It also defines a standardized benchmark that evaluates existing avatar fingerprinting systems and explores new methods based on Foundation Models such as DINOv2 and CLIP, while analyzing performance under generator and dataset shift. The authors find that identity‑related motion cues persist across synthetic avatars, yet current fingerprinting systems are highly sensitive to changes in synthesis pipelines and source domains.

By Laura Pedrouzo-Rodriguez, Luis F. Gomez, Ruben Tolosana, Ruben Vera-Rodriguez, Roberto Daza, Aythami Morales, Julian Fierrez
arXiv AI
Sep 12

A Survey of Threats Against Voice Authentication and Anti-Spoofing Systems

The paper reviews how voice authentication has evolved from handcrafted acoustic features to deep learning speaker embeddings, expanding its use in finance, smart devices, and law enforcement. It surveys modern threats—including data poisoning, adversarial, deepfake, and adversarial spoofing attacks—tracing their development alongside technological advances. For each attack type, the authors summarize methods, datasets, performance, and limitations, and organize the literature using accepted taxonomies to highlight emerging risks and open challenges.

By Kamel Kamel, Keshav Sood, Hridoy Sankar Dutta, Sunil Aryal
arXiv AI
Sep 12

Spectral Masking and Interpolation Attack (SMIA): A Black-box Adversarial Attack against Voice Authentication and Anti-Spoofing Systems

The paper introduces the Spectral Masking and Interpolation Attack (SMIA), a black‑box adversarial technique that subtly alters inaudible frequency regions of AI‑generated audio to fool voice authentication systems and their countermeasures. Experiments show SMIA achieves at least 82% success against combined verification and countermeasure systems, 97.5% against standalone speaker verification, and 100% against countermeasures, revealing a critical security gap. The authors argue that current static defenses are inadequate and call for dynamic, context‑aware defenses that can adapt to evolving threats.

By Kamel Kamel, Hridoy Sankar Dutta, Keshav Sood, Sunil Aryal