arXiv:2511. 05852v4 Announce Type: replace-cross Abstract: Knowledge editing (KE) offers a lightweight alternative to retraining for updating large language models (LLMs).
By Yinjie Cheng, Paul Youssef, Christin Seifert, J\"org Schl\"otterer, Zhixue Zhao
The paper investigates how tokenization can undermine post‑release guarantees that sensitive knowledge has been edited or unlearned from open‑weight large language models. By showing that alternative valid tokenizations can bypass localized modifications, the authors introduce Toketive, a reference‑free attack that detects modified knowledge and reconstructs pre‑edit responses using only the released model. Experiments on five LLMs, six datasets, and six editing techniques reveal that 38.6% of alternative tokenizations recover suppressed information, with Toketive achieving high detection and reconstruction accuracy.
By Manit Baser, Aditya Nawal, Dinil Mon Divakaran, Mohan Gurusamy
The paper investigates whether knowledge editing truly erases original facts from language models. Using a linear trace probe, the authors find that after editing a fact in GPT‑2‑XL, the original object remains highly decodable from hidden states across three different editing methods, even when the model behaves correctly on edited prompts. This suggests that editing suppresses rather than removes the original association in representational space.
By Priyansh Srivastava, Romit Chatterjee
The paper introduces ALOE, a method for knowledge editing that learns semantic addresses from paraphrases and hard negatives, aligns them with autoregressive hidden states, and embeds a gated low‑rank operator within a single MLP layer. This design allows the edited model to run in one forward pass without external retrievers or routers. Experiments on CounterFact, ZSRE, and KnowEdit show high efficacy (0.955–0.999) and locality (0.981–1.000) across 7–8B model families, with analyses indicating effective separation of edits and suppression of out‑of‑scope activation.
By Zeyan Li, Hu Xu, Jianfeng Xu
The paper introduces a new white‑box attack on large language models that builds on knowledge‑editing techniques. By retrieving associative knowledge from the model, the attack extends constraint removal to a whole thematic category rather than just a predefined set of prompts. Experiments across different architectures show that this method improves attack effectiveness while preserving overall model performance.
By Roman Maksimov, Vladimir Aletov, Vladimir Solodkin, Dmitry Bylinkin, Daniil Medyakov, Aleksandr Beznosikov
BARRIER (Bounded Activation Regions for Robust Information Erasure) is a method for machine unlearning that confines parameter updates to a controlled activation space, allowing stronger erasure of targeted concepts while limiting collateral damage to other representations. By employing interval arithmetic, it derives a closed‑form bound on worst‑case representation changes in protected regions, which serves as a knowledge‑preservation objective. The approach is architecture‑agnostic, compatible with existing erasure objectives, and empirically shows competitive performance in both classification and generative tasks, with enhanced robustness against adversarial recovery attacks.
By Jan Miksa, Patryk Krukowski, Przemys{\l}aw Spurek, Dawid Damian Rymarczyk, Marcin Sendera
arXiv:2608. 11660v1 Announce Type: cross Abstract: Large language models (LLMs) achieve remarkable performance across natural language tasks, yet they are trained on static corpora and their knowledge quickly becomes outdated in a fast-changing world.
By Tianci Liu, Zihan Dong, Tianchun Li, Yi-Chung Chen, Qiming Cao, Xingchen Wang, Shiyang Wang, Zichen Miao, Linjun Zhang, Haoyu Wang, Jing Gao
arXiv:2607. 20435v1 Announce Type: cross Abstract: Open-source LLMs (OSMs)arereaching near state-of-the-art performance, prompting prior works to trace the text they generate by embedding text watermarking algorithms directly into their weights.
By Luisa Scharff, Thibaud Gloaguen, Robin Staab, Martin Vechev
arXiv:2607. 20433v1 Announce Type: cross Abstract: While language models remain frozen at their training state, the world evolves continuously.
By Jea Kwon, Jiwon Kim, Dong-kyum Kim, Meeyoung Cha
arXiv:2605. 07961v2 Announce Type: replace Abstract: Federated fine-tuning (FFT) has emerged as a privacy-preserving paradigm for collaboratively adapting large language models (LLMs).
By Hanlin Cai, Kai Li, Houtianfu Wang, Haofan Dong, Yichen Li, Falko Dressler, Ozgur B. Akan
The paper introduces GLIME, a lifelong model editing framework that integrates knowledge editing with preference optimization to handle continual updates in large language models. GLIME employs replay-based editing and a gradient constraint to prevent overfitting to target prompts and preserve previously edited knowledge. Experiments demonstrate that GLIME enhances knowledge generalization while maintaining editing performance and overall model capabilities.
By Dahyun Jung, Suhyune Son, Heuiseok Lim
arXiv:2606. 18309v1 Announce Type: cross Abstract: Large Language Model (LLM) unlearning aims to remove undesirable knowledge or behaviors while preserving retained capabilities.
By Jingyuan Zhang, Yucheng Bai, Peixi Wen, Zhehao Huang, Zhengbao He, Hanling Tian, Xinwen Cheng, Haiyin Ran, Xiaolin Huang