The paper "Beyond EER: Multi-Dimensional Evaluation of Information Leakage in Speaker De-Identification" argues that evaluating speaker de-identification systems solely by Equal Error Rate (EER) is insufficient. It proposes a holistic framework using five complementary metrics—EER, soft biometric leakage score, cumulative match characteristic re-identification analysis, canonical correlation analysis with Procrustes embedding alignment, and intelligibility via word error rate and semantic similarity—to capture independent dimensions of information leakage. Experiments on five IARPA ARTS SDID systems show that these metrics reveal leakage that a single metric would miss.
By Seungmin Seo, Oleg Aulov, P. Jonathon Phillips, Kevin Mangold, Jonathan Eskin
arXiv:2607. 21820v1 Announce Type: cross Abstract: Audio deepfake detectors are trained to distinguish genuine speech from synthetic speech and often perform well on standard benchmarks.
By Daniyal Kabir Dar, Arun Ross
The paper introduces AVAPrintDB, a new public multi‑generator talking‑head avatar database designed for avatar fingerprinting, comprising data from two audiovisual corpora and three state‑of‑the‑art generators (GAGAvatar, LivePortrait, HunyuanPortrait). It also defines a standardized benchmark that evaluates existing avatar fingerprinting systems and explores new methods based on Foundation Models such as DINOv2 and CLIP, while analyzing performance under generator and dataset shift. The authors find that identity‑related motion cues persist across synthetic avatars, yet current fingerprinting systems are highly sensitive to changes in synthesis pipelines and source domains.
By Laura Pedrouzo-Rodriguez, Luis F. Gomez, Ruben Tolosana, Ruben Vera-Rodriguez, Roberto Daza, Aythami Morales, Julian Fierrez
arXiv:2606. 08678v1 Announce Type: cross Abstract: Sophisticated generative speech technology can undermined the reliability of voice biometrics.
By Anh-Tuan Dao, Driss Matrouf, Mickael Rouvier, Nicholas Evans
arXiv:2604. 01562v2 Announce Type: replace-cross Abstract: Voice cloning is often evaluated in terms of overall quality, but less is known about accent preservation and its perceptual consequences.
By Tianle Yang, Chengzhe Sun, Phil Rose, Siwei Lyu
The paper reviews how voice authentication has evolved from handcrafted acoustic features to deep learning speaker embeddings, expanding its use in finance, smart devices, and law enforcement. It surveys modern threats—including data poisoning, adversarial, deepfake, and adversarial spoofing attacks—tracing their development alongside technological advances. For each attack type, the authors summarize methods, datasets, performance, and limitations, and organize the literature using accepted taxonomies to highlight emerging risks and open challenges.
By Kamel Kamel, Keshav Sood, Hridoy Sankar Dutta, Sunil Aryal