arXiv AI By Haoyu Zhang, Shibo Zheng, Xiangchen Guan, Zhuoxi Wang, Zijian Xiao, Mohammad Zandsalimy, Shanu Sushmita

Borrowed Strength: Best-of-N Search over a Code EncodingBreaks Self-Check Jailbreak Defenses

Read the original on arXiv AI →

arXiv:2607. 26639v1 Announce Type: cross Abstract: A self-check defense asks the target model to assess a request before answering it; SAGE, the strongest published instance, reports an average 99% defense success rate.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

Hugging Face Trending Papers
Jul 29

Borrowed Strength: Best-of-N Search over a Code EncodingBreaks Self-Check Jailbreak Defenses

A self-check defense asks the target model to assess a request before answering it; SAGE, the strongest published instance, reports an average 99% defense success rate. We show it can be breached by composing two attacks that are individually harmless against it: an established code-completion encoding and an established best-of-N search, neither of which exceeds 4.

arXiv AI
6d ago

Depth, Not Breadth: Best-of-N Jailbreaking Beyond Surface Noise

The paper investigates how allocating a query budget to structural depth rather than surface variation improves jailbreak success against the SAGE self‑check defense. By using a best‑of‑N approach over a code‑completion encoding, the authors achieve 67%, 22%, and 15% success rates on three open‑weight targets—far exceeding the 4.7% and 3.0% rates of single‑draw encoding and character‑search methods. The study demonstrates that depth of encoding and breadth of variation independently undermine transform and gate defenses, and that repeated sampling can inflate perceived robustness.

By Haoyu Zhang, Hanwen Liu, Yang Chen, Shibo Zheng, Xiangchen Guan, Zhuoxi Wang, Zijian Xiao, Xiao Luo, Yi Feng, Haowen Xu, Mohammad Zandsalimy, Shanu Sushmita
arXiv AI
Aug 11

Attack Ensembles Expose a Safety-Utility Trade-off in Black-Box Guard Defenses Against Encoded VLM Jailbreaks

arXiv:2607. 26574v2 Announce Type: replace-cross Abstract: Safety classifiers ("guards") are the dominant black-box defense for vision-language models, yet a guard judges an input's surface form, not its meaning: a harmful request re-encoded as set theory, formal logic, a classical language, code, or text rendered inside an image slips past a guard that would block it in plain language - the decode gap.

By Haoyu Zhang, Zhuoxi Wang, Shibo Zheng, Yi Feng, Xiao Luo, Zijian Xiao, Haowen Xu, Xiangchen Guan, Mohammad Zandsalimy, Shanu Sushmita
arXiv Machine Learning
Jul 30

Recover, Decode, Reguard: Guard-Agnostic Defense Amplification againstEncoded VLM Jailbreaks

arXiv:2607. 26574v1 Announce Type: cross Abstract: Safety classifiers ("guards") are the dominant black-box defense for vision-language models, yet they judge an input's surface form, not its meaning: a harmful request re-encoded as set theory, formal logic, a rare language, code, or an image of text slips past a guard that would block it in plain language -- the decode gap.

By Haoyu Zhang, Zhuoxi Wang, Shibo Zheng, Zijian Xiao, Xiangchen Guan, Mohammad Zandsalimy, Shanu Sushmita
arXiv AI
Jun 2

Claudini: Autoresearch Discovers State-of-the-Art Adversarial Attack Algorithms for LLMs

arXiv:2603. 24511v2 Announce Type: replace-cross Abstract: We show that AI agents are capable of discovering novel algorithms for adversarial attacks against LLMs, advancing the state of the art on white-box jailbreaking and prompt injection evaluations.

By Alexander Panfilov, Peter Romov, Igor Shilov, Yves-Alexandre de Montjoye, Jonas Geiping, Maksym Andriushchenko
arXiv AI
Sep 25

Refusing Everything Looks Safe: Restoring the Benign Arm to Encoded-Prompt Evaluation

The paper critiques current encoded‑prompt safety benchmarks that focus only on harmful requests, showing that such tests can misrepresent a model’s safety. By evaluating the benign arm under the same encoding, the authors reveal a substantial drop in the harm gap—sometimes to zero—indicating that the encoding masks true safety deficiencies. Across multiple large models and training pipelines, they document that the encoding can either hide or falsely inflate safety metrics, and they identify twelve specific instrument defects that contribute to these misleading results.

By Haoyu Zhang, Haowen Xu, Xiao Luo, Hanwen Liu, Yang Chen, Zijian Xiao, Yi Feng, Xiangchen Guan, Mohammad Zandsalimy, Shanu Sushmita