arXiv AI By Berkay Ozcam, Irem Onen, Mehmet Fatih Amasyali, Emin Islam Tatli

Generating Attacks for LLMs with GFlowNets

Read the original on arXiv AI →

arXiv:2608. 10171v1 Announce Type: new Abstract: The rapid advancement of Large Language Models (LLMs) has facilitated their ubiquitous integration into various domains, leading to widespread adoption.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Machine Learning
Sep 1

Learning diverse attacks on large language models for robust red-teaming and safety tuning

arXiv:2405.18540v3 Announce Type: replace-cross Abstract: Red-teaming, or identifying prompts that elicit harmful responses, is a critical step in ensuring the safe and responsible deployment of larg...

By Seanie Lee, Minsu Kim, Lynn Cherif, David Dobre, Juho Lee, Sung Ju Hwang, Kenji Kawaguchi, Gauthier Gidel, Yoshua Bengio, Esmeralda S. Whitammer, Moksh Jain
arXiv Machine Learning
Aug 4

Quality-Diversity Red-Teaming: Automated Generation of High-Quality and Diverse Attackers for Large Language Models

arXiv:2506. 07121v2 Announce Type: replace Abstract: Ensuring the safety and robustness of large language models (LLMs) is a fundamental challenge and a critical prerequisite for the responsible deployment of artificial intelligence.

By Ren-Jian Wang, Ke Xue, Zeyu Qin, Ziniu Li, Sheng Tang, Hao-Tian Li, Shengcai Liu, Zhi Yu, Yuanpeng Tan, Chao Qian
arXiv AI
Sep 1

T-MAP: Red-Teaming LLM Agents with Trajectory-aware Evolutionary Search

The paper introduces T-MAP, a trajectory‑aware evolutionary search technique designed to red‑team large language model agents by exploiting vulnerabilities that arise during multi‑step tool execution. Unlike traditional methods that focus on harmful text, T‑MAP uses execution trajectories to generate adversarial prompts that bypass safety guardrails and achieve harmful objectives through actual tool interactions. Experiments across various Model Context Protocol environments show that T‑MAP outperforms baseline methods in attack realization rate and remains effective against advanced models such as GPT‑5.2, Gemini‑3‑Pro, Qwen3.5, and GLM‑5.

By Hyomin Lee, Sangwoo Park, Yumin Choi, Sohyun An, Seanie Lee, Sung Ju Hwang