arXiv AI

ARENA: Automated Red-Teaming for Large Audio Language Models

arXiv:2608. 15578v1 Announce Type: cross Abstract: Large audio-language models (LALMs) make it possible to interact with language models through speech, music, and environmental sound, but they also introduce a safety surface that is difficult to expose with text-only red-teaming.

arXiv Machine Learning
Aug 4

Quality-Diversity Red-Teaming: Automated Generation of High-Quality and Diverse Attackers for Large Language Models

arXiv:2506. 07121v2 Announce Type: replace Abstract: Ensuring the safety and robustness of large language models (LLMs) is a fundamental challenge and a critical prerequisite for the responsible deployment of artificial intelligence.

By Ren-Jian Wang, Ke Xue, Zeyu Qin, Ziniu Li, Sheng Tang, Hao-Tian Li, Shengcai Liu, Zhi Yu, Yuanpeng Tan, Chao Qian
arXiv Machine Learning
Sep 1

Learning diverse attacks on large language models for robust red-teaming and safety tuning

arXiv:2405.18540v3 Announce Type: replace-cross Abstract: Red-teaming, or identifying prompts that elicit harmful responses, is a critical step in ensuring the safe and responsible deployment of larg...

By Seanie Lee, Minsu Kim, Lynn Cherif, David Dobre, Juho Lee, Sung Ju Hwang, Kenji Kawaguchi, Gauthier Gidel, Yoshua Bengio, Esmeralda S. Whitammer, Moksh Jain
arXiv AI
Aug 20

`From Prompt to Perturbation': An Adaptive Framework for Voice-Based Jailbreaks on Audio LLMs

The paper introduces an adaptive jailbreak attack framework that evaluates both cascaded pipelines and end‑to‑end large audio‑language models (LALMs) under a unified setting. It employs a feedback‑guided mutation engine to automatically generate and refine jailbreak candidates across textual prompts and audio perturbations, thereby broadening attack diversity. Experiments on six audio‑based systems show that both paradigms remain highly vulnerable, with the framework achieving higher attack success rates than existing methods.

By Linghan Huang, Bo Li, Huaming Chen, Kim-Kwang Raymond Choo
arXiv AI
Aug 28

SpeechGym: An Audio-Native Gym for Training Voice Agents via Reinforcement Learning

SpeechGym is an audio‑native environment that lets two omni‑modal models converse entirely in native audio, eliminating external ASR/TTS and API boundaries while preserving the tasks, tools, and success checks of a standard text‑based agent benchmark. By training end‑to‑end, the framework addresses perceptual failures—such as misheard arguments that cascade into failed calls—and behavioural failures, both of which are automatically labeled for free. Using per‑turn process rewards to overcome reward sparsity, agents trained in SpeechGym transfer to an independent voice benchmark, doubling task success and improving efficiency in turns and tokens.

By Jiajun Fan, Jingyuan Li, Prashanth Gurunath Shivakumar, Jia-Hong Huang, Qi Luo, M. Maruf, Ivan Bulyko, Ge Liu, Roger Ren
arXiv Computation and Language
Sep 17

FRAUDSkill: Structured Frozen-Weight Skill Optimization for Audio Anti-Fraud Detection

FRAUDSkill is a structured frozen‑weight adaptation framework for audio anti‑fraud detection that keeps the underlying audio‑language model unchanged while optimizing external skill programs, route‑specific policies, and decision rules. It combines structured output control with validation‑guided multi‑path inference to produce protocol‑compliant predictions. On the TeleAntiFraud benchmark, FRAUDSkill achieves a 73.50% Macro‑F1 score, outperforming the shared frozen‑model baseline by 31.96% and reducing invalid outputs to 1.94%.

By Chengxian Hu, Zhiming Ma, Mingjun Pan, Yifan Wang, Shun Zhang, Qifan Wang, Zhilei Zhao, Yijin Zhou, Yuxi Zhao, Huiyuan Liu, Peidong Wang, Peng Chen
arXiv AI
Aug 18

JailbreakSkill: Scaling Automated Red-Teaming with Reusable and Ever-Evolving Skills

arXiv:2608. 16465v1 Announce Type: new Abstract: Automated red-teaming has produced a growing collection of attack strategies, yet they typically remain scattered across prompts and workflows, making them difficult to systematically integrate, reuse, and improve at scale.

By Xiaoyu Wen, Jiajia Li, Zhida He, Peng Yu, Chenxu Wang, Han Qi, Ziyuan Zhou, Cheng Jin, Ying Wen, Xingcheng Xu, Shuyue Hu, Tianhang Zheng, Chaochao Lu, Qiaosheng Zhang
arXiv AI
Aug 28

RedEvoAgent: Automatic Red-Teaming Agent with Experience-Driven Skill Evolution

RedEvoAgent is a black-box red‑teaming agent that transforms cross‑case attack trajectories into concise, human‑readable attack skills. It evolves these skills by profiling tool effectiveness, attributing tool credit, and applying a validation ratchet to keep only improvements. Experiments demonstrate that RedEvoAgent outperforms fixed and agentic baselines, enhances tool efficiency, and transfers across attacker models and target execution harnesses.

By Junjie Zhang, Hui Liu, Kecheng Chen, Xianbo Mo, Changsheng Chen, Haoliang Li
arXiv AI
Sep 1

T-MAP: Red-Teaming LLM Agents with Trajectory-aware Evolutionary Search

The paper introduces T-MAP, a trajectory‑aware evolutionary search technique designed to red‑team large language model agents by exploiting vulnerabilities that arise during multi‑step tool execution. Unlike traditional methods that focus on harmful text, T‑MAP uses execution trajectories to generate adversarial prompts that bypass safety guardrails and achieve harmful objectives through actual tool interactions. Experiments across various Model Context Protocol environments show that T‑MAP outperforms baseline methods in attack realization rate and remains effective against advanced models such as GPT‑5.2, Gemini‑3‑Pro, Qwen3.5, and GLM‑5.

By Hyomin Lee, Sangwoo Park, Yumin Choi, Sohyun An, Seanie Lee, Sung Ju Hwang