arXiv Machine Learning

Don't Walk the Line: Boundary Guidance for Filtered Generation

arXiv:2510. 11834v3 Announce Type: replace Abstract: Generative models are increasingly paired with safety classifiers that filter harmful or undesirable outputs.

arXiv AI
Jun 4

REFLECTOR: Internalizing Step-wise Reflection against Indirect Jailbreak

arXiv:2605. 20654v2 Announce Type: replace-cross Abstract: While Large Language Models (LLMs) demonstrate remarkable capabilities, they remain susceptible to sophisticated, multi-step jailbreak attacks that circumvent conventional surface-level safety alignment by exploiting the internal generation process.

By Jiachen Ma, Jiawen Zhang, Xiangtian Li, Bo Zou, Chaochao Lu, Chao Yang