arXiv:2606. 05678v1 Announce Type: cross Abstract: Automatic speech recognition (ASR) systems have become widely used for multilingual speech-to-text transcription.
By Yifan Liao, Zongmin Zhang, Zhen Sun, Yuhui Sun, Xinhu Zheng, Xinlei He
arXiv:2607. 14753v1 Announce Type: cross Abstract: Recent advances in text-to-speech and voice cloning make high-quality spoofing inexpensive and scalable, threatening voice authentication systems, especially automatic speaker verification (ASV).
By Sofya Savelyeva, Mariia Perunova, Evgeny Kushnir, Artem Dvirniak, Dmitrii Korzh, Oleg Y. Rogov
arXiv:2606. 16837v1 Announce Type: cross Abstract: Spoofed speech detection is increasingly challenged by realistic synthesis, voice conversion, and replay attacks, with cross-dataset generalization remaining a major limitation.
By Mahtab Masoudi Nezhad, Nima Karimian
arXiv:2606. 08678v1 Announce Type: cross Abstract: Sophisticated generative speech technology can undermined the reliability of voice biometrics.
By Anh-Tuan Dao, Driss Matrouf, Mickael Rouvier, Nicholas Evans
arXiv:2606. 31411v1 Announce Type: cross Abstract: Rapid advancements in generative speech technology have compromised the reliability of voice biometrics.
By Anh-Tuan Dao, Driss Matrouf, Mickael Rouvier, Nicholas Evans
arXiv:2608. 10405v1 Announce Type: cross Abstract: Many studies have shown that specially crafted inputs can induce large language models (LLMs) to generate excessively long outputs, resulting in significant computational overhead and resource consumption.
By Shuozhe Cheng, Kunlan Xiang, Mingxuan Li, Ji Zhang, Dongxiao Liu, Wenbo Jiang
arXiv:2607. 16870v1 Announce Type: cross Abstract: End-to-end speech language models increasingly represent user speech with speech tokens rather than relying exclusively on cascaded ASR--LLM--TTS pipelines.
By Ye Lu, Yihan Yan, Zhaoyang Zhang, Zhitao Ou, Runze Liu, Li Liu, Shen Wang
arXiv:2607. 11706v1 Announce Type: cross Abstract: Modern LLM-driven text-to-speech (TTS) and voice conversion (VC) systems produce synthetic speech that differs from the generators represented in many legacy spoofing benchmarks.
By Aastha Sharma, Guangjing Wang
arXiv:2603. 22590v2 Announce Type: replace Abstract: With the increasing deployment of automated and agentic systems, ensuring the adversarial robustness of automatic speech recognition (ASR) models has become highly relevant.
By Mat\'ias Pizarro, Raghavan Narasimhan, Jonas Killian, Asja Fischer
arXiv:2602. 02557v2 Announce Type: replace-cross Abstract: Recent advances in end-to-end trained omni-models have substantially improved audio capabilities by strengthening text-audio modality alignment.
By Yupeng Chen, Junchi Yu, Aoxi Liu, Baoyuan Wu, Philip Torr, Adel Bibi
arXiv:2606. 29544v1 Announce Type: cross Abstract: We present Proteus, a framework developed at Resemble AI for automated robustness testing of our audio deepfake detection system.
By Nicolas M. M\"uller, Aditya Tirumala Bukkapatnam, Zohaib Ahmed
arXiv:2606. 10246v1 Announce Type: cross Abstract: Maliciously-created fake speech, including deepfaked and spoofed audio, is proliferating at an alarming rate, and detection models are racing to stay ahead of the curve.
By Ashley R. Keaton, Zahra Khanjani, Christine Mallinson, Vandana P. Janeja