arXiv AI By Yongbin Huang, Xihao Xie, Jia Zhang

Backdoor Attacks on Speech Emotion Recognition via TTS-Generated Poisoning

Read the original on arXiv AI →

The paper investigates poisoning-based backdoor attacks on Speech Emotion Recognition (SER) systems that use self‑supervised acoustic representations. It introduces a stealthy, low‑energy acoustic trigger that can be embedded imperceptibly into both natural and synthetic speech, enabling scalable poisoning. Experiments show high attack success rates with low poisoning ratios, cross‑model transferability, and a particular vulnerability of self‑supervised representations, highlighting the lowered barrier to effective backdoor attacks via TTS technology.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Aug 20

`From Prompt to Perturbation': An Adaptive Framework for Voice-Based Jailbreaks on Audio LLMs

The paper introduces an adaptive jailbreak attack framework that evaluates both cascaded pipelines and end‑to‑end large audio‑language models (LALMs) under a unified setting. It employs a feedback‑guided mutation engine to automatically generate and refine jailbreak candidates across textual prompts and audio perturbations, thereby broadening attack diversity. Experiments on six audio‑based systems show that both paradigms remain highly vulnerable, with the framework achieving higher attack success rates than existing methods.

By Linghan Huang, Bo Li, Huaming Chen, Kim-Kwang Raymond Choo