arXiv:2511. 05852v4 Announce Type: replace-cross Abstract: Knowledge editing (KE) offers a lightweight alternative to retraining for updating large language models (LLMs).
By Yinjie Cheng, Paul Youssef, Christin Seifert, J\"org Schl\"otterer, Zhixue Zhao
The paper investigates how tokenization can undermine post‑release guarantees that sensitive knowledge has been edited or unlearned from open‑weight large language models. By showing that alternative valid tokenizations can bypass localized modifications, the authors introduce Toketive, a reference‑free attack that detects modified knowledge and reconstructs pre‑edit responses using only the released model. Experiments on five LLMs, six datasets, and six editing techniques reveal that 38.6% of alternative tokenizations recover suppressed information, with Toketive achieving high detection and reconstruction accuracy.
By Manit Baser, Aditya Nawal, Dinil Mon Divakaran, Mohan Gurusamy
The paper investigates whether knowledge editing truly erases original facts from language models. Using a linear trace probe, the authors find that after editing a fact in GPT‑2‑XL, the original object remains highly decodable from hidden states across three different editing methods, even when the model behaves correctly on edited prompts. This suggests that editing suppresses rather than removes the original association in representational space.
By Priyansh Srivastava, Romit Chatterjee
The paper introduces ALOE, a method for knowledge editing that learns semantic addresses from paraphrases and hard negatives, aligns them with autoregressive hidden states, and embeds a gated low‑rank operator within a single MLP layer. This design allows the edited model to run in one forward pass without external retrievers or routers. Experiments on CounterFact, ZSRE, and KnowEdit show high efficacy (0.955–0.999) and locality (0.981–1.000) across 7–8B model families, with analyses indicating effective separation of edits and suppression of out‑of‑scope activation.
By Zeyan Li, Hu Xu, Jianfeng Xu
The paper introduces a new white‑box attack on large language models that builds on knowledge‑editing techniques. By retrieving associative knowledge from the model, the attack extends constraint removal to a whole thematic category rather than just a predefined set of prompts. Experiments across different architectures show that this method improves attack effectiveness while preserving overall model performance.
By Roman Maksimov, Vladimir Aletov, Vladimir Solodkin, Dmitry Bylinkin, Daniil Medyakov, Aleksandr Beznosikov
BARRIER (Bounded Activation Regions for Robust Information Erasure) is a method for machine unlearning that confines parameter updates to a controlled activation space, allowing stronger erasure of targeted concepts while limiting collateral damage to other representations. By employing interval arithmetic, it derives a closed‑form bound on worst‑case representation changes in protected regions, which serves as a knowledge‑preservation objective. The approach is architecture‑agnostic, compatible with existing erasure objectives, and empirically shows competitive performance in both classification and generative tasks, with enhanced robustness against adversarial recovery attacks.
By Jan Miksa, Patryk Krukowski, Przemys{\l}aw Spurek, Dawid Damian Rymarczyk, Marcin Sendera