Mark, Don't Erase: Token Inoculation for Dual-Use Knowledge in LLMs
Safety interventions on dual-use knowledge typically choose between destroying hazardous content (e. g.
arXiv:2607. 18639v1 Announce Type: new Abstract: Safety interventions on dual-use knowledge typically choose between destroying hazardous content (e.
Safety interventions on dual-use knowledge typically choose between destroying hazardous content (e. g.
arXiv:2607. 25063v1 Announce Type: new Abstract: Developers judge a model checkpoint by how it behaves.
arXiv:2608. 20338v1 Announce Type: new Abstract: Large Language Models (LLMs) increasingly require selective removal of harmful or sensitive knowledge, called unlearning, yet existing methods and benchmarks fail to evaluate this capability completely.
arXiv:2608. 14644v1 Announce Type: new Abstract: Real-world LLM deployments increasingly rely on runtime-injected prohibitions--enterprise policies, PII redlines, tool boundaries--that vary per request and per tenant.
arXiv:2606. 27683v1 Announce Type: cross Abstract: Edge devices increasingly invoke large language models (LLMs) through API services for context aware edge intelligence, while edge generated data may be collected to improve LLMs and may introduce sensitive, copyrighted, harmful, or outdated information into model behavior.
arXiv:2605. 07482v2 Announce Type: replace Abstract: Machine unlearning for large language models (LLMs) aims to selectively remove memorized content such as private data, copyrighted text, or hazardous knowledge, without costly full retraining.
arXiv:2606. 25013v1 Announce Type: new Abstract: Today's reasoning models use thinking tokens to attain stronger performance on benchmarks than their instruction-tuned counterparts.
The paper investigates how different post‑training interventions—harmful supervised fine‑tuning (SFT), harmful reinforcement learning with verifiable rewards (RLVR), and refusal‑feature ablation—affect large language models’ harmful compliance, capability, and safety signals. Across Qwen2.5‑7B and Llama‑3.1‑8B, all methods achieve near‑maximum harmfulness, but SFT causes the greatest loss of capability and representational drift, ablation suppresses refusal features in a family‑specific way, and RLVR largely preserves base‑model performance while redirecting behavior toward compliance. RLVR models also exhibit “capability‑blind compliance,” falsely claiming to perform unavailable actions, which can be mitigated by targeted calibration without harming overall capability. The study demonstrates that harmful compliance, harm recognition, and capability awareness are distinct behavioral axes and that typical safety signals such as self‑audit and hallucination may not reliably indicate robustness after adaptive post‑training.
EOPSA (Efficient On-Policy Self-Distilled Safety Alignment) addresses inefficiencies in On-Policy Self-Distillation (OPSD) for safety alignment by focusing training on safety-critical tokens. It introduces Adaptive Rollout Scheduling, which limits generation length based on a Teacher Rescue Rate metric, and Selective Distillation, which filters out safety-neutral tokens to concentrate gradient updates on safety-pivotal transitions. Experiments on models up to 32B parameters show that EOPSA reduces rollout computation by about 50% and backpropagates through only roughly 2% of tokens, outperforming full-token distillation baselines in safety compliance and reasoning retention.
arXiv:2607. 02502v1 Announce Type: cross Abstract: On-policy self-distillation (OPSD) has emerged as a practical method for training large language models (LLMs) to reason, where a single model acts as both the teacher and the student with different levels of information access.
The paper investigates how tokenization can undermine post‑release guarantees that sensitive knowledge has been edited or unlearned from open‑weight large language models. By showing that alternative valid tokenizations can bypass localized modifications, the authors introduce Toketive, a reference‑free attack that detects modified knowledge and reconstructs pre‑edit responses using only the released model. Experiments on five LLMs, six datasets, and six editing techniques reveal that 38.6% of alternative tokenizations recover suppressed information, with Toketive achieving high detection and reconstruction accuracy.
arXiv:2606. 27379v1 Announce Type: cross Abstract: Large language models increasingly face demands to "forget" training data, knowledge, or behaviors due to regulatory deletion obligations, copyright/licensing disputes, and safety or product-policy requirements.