The paper "Beyond EER: Multi-Dimensional Evaluation of Information Leakage in Speaker De-Identification" argues that evaluating speaker de-identification systems solely by Equal Error Rate (EER) is insufficient. It proposes a holistic framework using five complementary metrics—EER, soft biometric leakage score, cumulative match characteristic re-identification analysis, canonical correlation analysis with Procrustes embedding alignment, and intelligibility via word error rate and semantic similarity—to capture independent dimensions of information leakage. Experiments on five IARPA ARTS SDID systems show that these metrics reveal leakage that a single metric would miss.
By Seungmin Seo, Oleg Aulov, P. Jonathon Phillips, Kevin Mangold, Jonathan Eskin
arXiv:2607. 21820v1 Announce Type: cross Abstract: Audio deepfake detectors are trained to distinguish genuine speech from synthetic speech and often perform well on standard benchmarks.
By Daniyal Kabir Dar, Arun Ross
The paper introduces AVAPrintDB, a new public multi‑generator talking‑head avatar database designed for avatar fingerprinting, comprising data from two audiovisual corpora and three state‑of‑the‑art generators (GAGAvatar, LivePortrait, HunyuanPortrait). It also defines a standardized benchmark that evaluates existing avatar fingerprinting systems and explores new methods based on Foundation Models such as DINOv2 and CLIP, while analyzing performance under generator and dataset shift. The authors find that identity‑related motion cues persist across synthetic avatars, yet current fingerprinting systems are highly sensitive to changes in synthesis pipelines and source domains.
By Laura Pedrouzo-Rodriguez, Luis F. Gomez, Ruben Tolosana, Ruben Vera-Rodriguez, Roberto Daza, Aythami Morales, Julian Fierrez
arXiv:2606. 08678v1 Announce Type: cross Abstract: Sophisticated generative speech technology can undermined the reliability of voice biometrics.
By Anh-Tuan Dao, Driss Matrouf, Mickael Rouvier, Nicholas Evans
arXiv:2604. 01562v2 Announce Type: replace-cross Abstract: Voice cloning is often evaluated in terms of overall quality, but less is known about accent preservation and its perceptual consequences.
By Tianle Yang, Chengzhe Sun, Phil Rose, Siwei Lyu
The paper reviews how voice authentication has evolved from handcrafted acoustic features to deep learning speaker embeddings, expanding its use in finance, smart devices, and law enforcement. It surveys modern threats—including data poisoning, adversarial, deepfake, and adversarial spoofing attacks—tracing their development alongside technological advances. For each attack type, the authors summarize methods, datasets, performance, and limitations, and organize the literature using accepted taxonomies to highlight emerging risks and open challenges.
By Kamel Kamel, Keshav Sood, Hridoy Sankar Dutta, Sunil Aryal
arXiv:2606. 28048v1 Announce Type: cross Abstract: Insurance fraud remains costly and operationally difficult, particularly in call-centre workflows where many customer interactions begin at FNOL.
By Muhammad Shakeel Akram, Amal Htait, Abdul Hamid Sadka, Emma Meisingseth, Karishma Jaitly
The paper introduces the Spectral Masking and Interpolation Attack (SMIA), a black‑box adversarial technique that subtly alters inaudible frequency regions of AI‑generated audio to fool voice authentication systems and their countermeasures. Experiments show SMIA achieves at least 82% success against combined verification and countermeasure systems, 97.5% against standalone speaker verification, and 100% against countermeasures, revealing a critical security gap. The authors argue that current static defenses are inadequate and call for dynamic, context‑aware defenses that can adapt to evolving threats.
By Kamel Kamel, Hridoy Sankar Dutta, Keshav Sood, Sunil Aryal
arXiv:2607. 03985v1 Announce Type: cross Abstract: Advanced neural technologies in speech synthesis and voice conversion (VC) have introduced severe risks to personal privacy, necessitating robust Speaker Anonymization Systems (SAS).
By Meiying Melissa Chen, Anastasia Kuznetsova, Zhenyu Wang, Zhiyao Duan
arXiv:2603. 14033v2 Announce Type: replace-cross Abstract: Audio anti-spoofing systems are typically trained to assign one authenticity label to an entire speech utterance.
By Shree Harsha Bokkahalli Satish, Harm Lameris, Joakim Gustafson, \'Eva Sz\'ekely
arXiv:2608. 15411v1 Announce Type: new Abstract: The ability of artificial intelligence (AI) models to generate highly realistic human voices has advanced rapidly.
By Chengzhe Sun, Tianle Yang, Siwei Lyu
arXiv:2608.30951v1 Announce Type: new
Abstract: The rapid development of generative portrait models has raised growing concerns about privacy leakage and identity misuse. In particular, audio-driven...
By Rui-Qing Sun, Chen-Hao Cui, Hui-Yang Zhao, Tian Lan, Zhijing Wu, Xian-Ling Mao