HE-Guardrail is a framework that applies homomorphic encryption to enforce guardrails against jailbreak attacks during encrypted large language model inference. It evaluates guardrail mechanisms—Llama Guard, JBShield, and GradSafe—directly on encrypted data, deciding whether to return the model’s response to the client. The approach preserves confidentiality while closely matching the decisions of plaintext guardrails, offering varied security-efficiency-utility trade‑offs.
By Byeongseo Min, Yongwoo Lee, Young-Sik Kim, Yongjune Kim
arXiv:2510. 01529v3 Announce Type: replace Abstract: Ball et al.
By Jaiden Fairoze, Sanjam Garg, Keewoo Lee, Mingyuan Wang
arXiv:2508. 10031v2 Announce Type: replace-cross Abstract: While Large Language Models (LLMs) have shown significant advancements in performance, various jailbreak attacks have posed growing safety and ethical risks.
By Jinhwa Kim, Ian G. Harris
arXiv:2609.05850v1 Announce Type: cross
Abstract: Despite the significant efforts devoted to aligning large language models (LLMs) with human values and ensuring safe deployment, recent work has reve...
By Quoc Viet Vo, Trung Le, Damith C. Ranasinghe, Ehsan Abbasnejad
arXiv:2609.26185v1 Announce Type: cross
Abstract: Large Language Models (LLMs) demonstrate impressive capabilities across many applications but remain vulnerable to jailbreak attacks, which elicit ha...
By Doniyorkhon Obidov, Honggang Yu, Xiaolong Guo, Kaichen Yang
arXiv:2606. 11817v1 Announce Type: cross Abstract: Large Language Models (LLMs) are increasingly used for code generation, raising concerns that they may be misused to produce malicious code.
By Yitong Zhang, Shiteng Lu, Jia Li