arXiv:2607. 01702v1 Announce Type: cross Abstract: Recently, speech classification methods have gained widespread adoption in intelligent gadgets.
By Yueming Huang, Wenhan Yao, Fen Xiao, Xiarun Chen, Weiping Wen
Recently, speech classification methods have gained widespread adoption in intelligent gadgets. Current study indicates that backdoor attacks provide a substantial security concern to these models, underscoring the pressing necessity to investigate additional potential attack techniques to expose and prevent such risks.
arXiv:2607. 15697v1 Announce Type: cross Abstract: Backdoor attacks pose a critical threat to neural network models, allowing attackers to implant a backdoor during the training phase by manipulating a small portion of the training data.
By Jinwen Xin, Xixiang Lv
arXiv:2607. 15724v1 Announce Type: cross Abstract: With the rapid development of deep learning, its vulnerability has gradually emerged in recent years.
By Jinwen Xin, Xixiang Lyu, Jing Ma
arXiv:2606. 05678v1 Announce Type: cross Abstract: Automatic speech recognition (ASR) systems have become widely used for multilingual speech-to-text transcription.
By Yifan Liao, Zongmin Zhang, Zhen Sun, Yuhui Sun, Xinhu Zheng, Xinlei He
The paper introduces an adaptive jailbreak attack framework that evaluates both cascaded pipelines and end‑to‑end large audio‑language models (LALMs) under a unified setting. It employs a feedback‑guided mutation engine to automatically generate and refine jailbreak candidates across textual prompts and audio perturbations, thereby broadening attack diversity. Experiments on six audio‑based systems show that both paradigms remain highly vulnerable, with the framework achieving higher attack success rates than existing methods.
By Linghan Huang, Bo Li, Huaming Chen, Kim-Kwang Raymond Choo
arXiv:2607. 01729v1 Announce Type: new Abstract: Deep learning models for speech classification are vulnerable to backdoor attacks, where malicious triggers cause misclassification at inference time.
By Yueming Huang, Wenhan Yao, Fen Xiao, Xiarun Chen, Weiping Wen
The paper introduces the Spectral Masking and Interpolation Attack (SMIA), a black‑box adversarial technique that subtly alters inaudible frequency regions of AI‑generated audio to fool voice authentication systems and their countermeasures. Experiments show SMIA achieves at least 82% success against combined verification and countermeasure systems, 97.5% against standalone speaker verification, and 100% against countermeasures, revealing a critical security gap. The authors argue that current static defenses are inadequate and call for dynamic, context‑aware defenses that can adapt to evolving threats.
By Kamel Kamel, Hridoy Sankar Dutta, Keshav Sood, Sunil Aryal
The paper reviews how voice authentication has evolved from handcrafted acoustic features to deep learning speaker embeddings, expanding its use in finance, smart devices, and law enforcement. It surveys modern threats—including data poisoning, adversarial, deepfake, and adversarial spoofing attacks—tracing their development alongside technological advances. For each attack type, the authors summarize methods, datasets, performance, and limitations, and organize the literature using accepted taxonomies to highlight emerging risks and open challenges.
By Kamel Kamel, Keshav Sood, Hridoy Sankar Dutta, Sunil Aryal
arXiv:2606. 06833v1 Announce Type: cross Abstract: Automatic Speech Recognition (ASR) systems operating in real-time settings must process acoustic input under strict temporal constraints, where transcription decisions are inherently made on incomplete information.
By Jiani Xie, Andrew C. Cullen, Paul Montague, Benjamin I. P. Rubinstein
arXiv:2607. 16870v1 Announce Type: cross Abstract: End-to-end speech language models increasingly represent user speech with speech tokens rather than relying exclusively on cascaded ASR--LLM--TTS pipelines.
By Ye Lu, Yihan Yan, Zhaoyang Zhang, Zhitao Ou, Runze Liu, Li Liu, Shen Wang
arXiv:2608. 10405v1 Announce Type: cross Abstract: Many studies have shown that specially crafted inputs can induce large language models (LLMs) to generate excessively long outputs, resulting in significant computational overhead and resource consumption.
By Shuozhe Cheng, Kunlan Xiang, Mingxuan Li, Ji Zhang, Dongxiao Liu, Wenbo Jiang