The paper investigates safety risks in model merging, showing that even when all constituent models are individually safety‑aligned, merging can expose a jailbreak vulnerability rooted in the pretrained foundation model. It introduces Basin‑Aware Jailbreak (BAJ), a min–max optimization method that generates adversarial suffixes transferable across merged models sharing the same backbone, without needing the exact merging coefficients or checkpoints. Experiments demonstrate BAJ’s high transfer success rates across diverse backbones and merging settings, and its resilience against existing defenses.
By Yu Zhe, Yixin Tan, Junhao Wei, Wang Chen
arXiv:2609.26185v1 Announce Type: cross
Abstract: Large Language Models (LLMs) demonstrate impressive capabilities across many applications but remain vulnerable to jailbreak attacks, which elicit ha...
By Doniyorkhon Obidov, Honggang Yu, Xiaolong Guo, Kaichen Yang
arXiv:2607. 07903v1 Announce Type: cross Abstract: Large language models (LLMs) exhibit remarkable capabilities but remain highly vulnerable to adversarial prompts and jailbreak attacks.
By Anupam Wagle, Ifrat Ikhtear Uddin, Chaowei Zhang, Longwei Wang
arXiv:2510.17904v3 Announce Type: replace-cross
Abstract: Large Language Models (LLMs) are widely used because they process structures, syntax and code well, but this same ability also makes them par...
By Amirkia Rafiei Oskooei, Mehmet S. Aktas
arXiv:2508. 10029v3 Announce Type: replace-cross Abstract: Safety-aligned large language models can still be manipulated through white-box interventions that modify their internal representations.
By Wenpeng Xing, Bohan Yang, Mohan Li, Chunqiang Hu, Haitao Xu, Ningyu Zhang, Bo Lin, Meng Han
arXiv:2607. 01859v1 Announce Type: new Abstract: Safety training for large language models (LLMs) is conducted predominantly in English, leaving uncertain how well safety mechanisms generalize to low-resource languages and mixed-language code-switching.
By Joshua Adrian Cahyono