arXiv:2606. 08678v1 Announce Type: cross Abstract: Sophisticated generative speech technology can undermined the reliability of voice biometrics.
By Anh-Tuan Dao, Driss Matrouf, Mickael Rouvier, Nicholas Evans
arXiv:2606. 16532v1 Announce Type: cross Abstract: Audio deepfake detectors often fail to generalize across speakers, as they learn speaker-identity features rather than synthesis artifacts, known as implicit identity leakage.
By Zhuodong Liu, Hugen Lv, Xiangyu Li, Chunhong Yuan
arXiv:2606. 08843v1 Announce Type: cross Abstract: We present a voice conversion (VC) framework that utilizes K-Nearest Neighbors (KNN) retrieval over WavLM representations to align non-parallel source and target speech, constructing synthetic training pairs for supervised learning.
By Moshe Mandel, Shlomo E. Chazan
arXiv:2601. 09239v5 Announce Type: replace-cross Abstract: Speech tokenizers are a key building block of fully discrete Speech LLMs.
By Hanlin Zhang, Daxin Tan, Dehua Tao, Xiao Chen, Haochen Tan, Yunhe Li, Yuchen Cao, Linqi Song
arXiv:2607. 03134v1 Announce Type: cross Abstract: Recent research expands beyond binary anti-spoofing with the emergence of Source Tracing, the task of identifying the specific generative origins of synthetic speech.
By Santiago Rubio, Antonio Almud\'evar, Antonio Miguel, Eduardo Lleida, Alfonso Ortega
arXiv:2606. 31664v1 Announce Type: cross Abstract: Performance in face and speaker verification is largely driven by margin-penalty softmax losses such as CosFace and ArcFace.
By Dimitrios Koutsianos, Ladislav Mo\v{s}ner, Yannis Panagakis, Themos Stafylakis