arXiv Computation and Language By Anjila Budathoki, Manish Dhakal, Benjamin M. Ampel, Yi Ding

Understanding the Role of Prompt Template in Knowledge Distillation for Safety Alignment

Read the original on arXiv Computation and Language →

The study investigates how prompt template choices during Knowledge Distillation (KD) affect safety alignment in language models. It finds that using chat templates during KD degrades safety alignment, making models more compliant with harmful queries, while non-chat templates better preserve the base model’s internal representations. These effects are observed across LLaMA, Gemma, and Qwen families on multiple safety benchmarks.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Computation and Language.

arXiv Computation and Language
Sep 3

SHARD: Safe and Helpful Alignment via Self-Reframing Distillation

SHARD is a self‑reframing distillation technique designed to enhance the safe‑helpfulness of large language models. It rewrites sensitive prompts to reveal benign intent, reframes the model’s original responses into safer, more helpful versions, and then fine‑tunes the model on these self‑reframed outputs. Experiments on DNA and the English subset of LINGUASAFE show that SHARD improves helpfulness across various model families while maintaining safety, and performs competitively with distillation from larger teacher models.

By Viswonathan Manoranjan, Amogh Gupta, Anvesh Rao Vijjini, Thomas Hofweber, Snigdha Chaturvedi
Hugging Face Trending Papers
6d ago

EOPSA: Efficient On-Policy Self-Distilled Safety Alignment

EOPSA (Efficient On-Policy Self-Distilled Safety Alignment) addresses inefficiencies in On-Policy Self-Distillation (OPSD) for safety alignment by focusing training on safety-critical tokens. It introduces Adaptive Rollout Scheduling, which limits generation length based on a Teacher Rescue Rate metric, and Selective Distillation, which filters out safety-neutral tokens to concentrate gradient updates on safety-pivotal transitions. Experiments on models up to 32B parameters show that EOPSA reduces rollout computation by about 50% and backpropagates through only roughly 2% of tokens, outperforming full-token distillation baselines in safety compliance and reasoning retention.