Breaking Refusal in the First Half: A Mechanistic Study of the Prefill Jailbreak
arXiv:2607. 14147v1 Announce Type: cross Abstract: Aligned language models refuse harmful requests, but a one-line prefill ("Sure, here is") strips the refusal.
arXiv:2606. 14388v1 Announce Type: new Abstract: Interventions designed to modify a particular behavior in LLMs, such as refusal or sycophancy, often produce unintended changes in other behaviors.
arXiv:2607. 14147v1 Announce Type: cross Abstract: Aligned language models refuse harmful requests, but a one-line prefill ("Sure, here is") strips the refusal.
arXiv:2606. 13720v1 Announce Type: new Abstract: Arditi et al.
arXiv:2608. 16177v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly deployed as agents that operate equipment, execute instructions, and act inside institutional hierarchies, raising a question social psychology answered for humans six decades ago: how far will an agent escalate a harmful action when a legitimate authority insists?
arXiv:2605. 00123v3 Announce Type: replace Abstract: Safety trained large language models (LLMs) can often be induced to answer harmful requests through jailbreak prompts.
arXiv:2608. 03201v1 Announce Type: new Abstract: Safety guards are widely used to filter harmful content and are typically trained via supervised fine-tuning on labeled prompt-response pairs.
arXiv:2509. 13450v3 Announce Type: replace Abstract: We introduce SteeringSafety, a benchmark for evaluating representation steering methods across nine safety perspectives spanning 18 datasets.
arXiv:2607. 12792v1 Announce Type: cross Abstract: Jailbreak-robustness research typically evaluates safety through generated responses using an LLM-as-judge approach.
arXiv:2607. 05355v1 Announce Type: cross Abstract: Attribution scores increasingly identify which neuron rows of a language model matter for applications such as pruning, interpretability, and editing for safety, yet whether they identify causally important rows is rarely tested directly.
arXiv:2608. 10703v1 Announce Type: cross Abstract: Large language models (LLMs) increasingly act in interactive settings where their behavioral styles affect user experience, safety, and downstream decision making.
arXiv:2606. 08044v1 Announce Type: cross Abstract: Large Language Model (LLM) safety has often been evaluated at the behavior level, which provides limited evidence of internal robustness, as these evaluations target outputs rather than representation-level vulnerability under intervention.
arXiv:2606. 08682v1 Announce Type: cross Abstract: Activation steering has emerged as a popular inference-time technique for modulating the behavior of large language models (LLMs).
arXiv:2606. 05614v1 Announce Type: new Abstract: Large language models (LLMs) are rigorously aligned to refuse harmful requests, a process that inherently cultivates a latent capacity to evaluate and recognize unsafe content.