arXiv Machine Learning By Sarah Ball, Andreas Haupt

Don't Walk the Line: Boundary Guidance for Filtered Generation

Read the original on arXiv Machine Learning →

arXiv:2510. 11834v3 Announce Type: replace Abstract: Generative models are increasingly paired with safety classifiers that filter harmful or undesirable outputs.

Summary generated by The Flow from the publisher's feed. The full article lives at arXiv Machine Learning.

arXiv AI
Jun 4

REFLECTOR: Internalizing Step-wise Reflection against Indirect Jailbreak

arXiv:2605. 20654v2 Announce Type: replace-cross Abstract: While Large Language Models (LLMs) demonstrate remarkable capabilities, they remain susceptible to sophisticated, multi-step jailbreak attacks that circumvent conventional surface-level safety alignment by exploiting the internal generation process.

By Jiachen Ma, Jiawen Zhang, Xiangtian Li, Bo Zou, Chaochao Lu, Chao Yang