arXiv Computation and Language By Guoshenghui Zhao, Huawei Lin, Weijie Zhao

RapidUn: Influence-Driven Parameter Reweighting for Efficient Large Language Model Unlearning

Read the original on arXiv Computation and Language →

RapidUn is a parameter reweighting framework that uses influence estimates to guide LoRA-only updates for efficient unlearning of targeted behaviors in large language models. It operates in a practical PEFT setting with a small forget set and limited retain buffer, converting cross-sample influence into fixed sample-specific weights for weighted LoRA unlearning. Experiments on Llama‑3‑8B with Dolly‑15k and Alpaca‑57k datasets show RapidUn achieves lower trigger ASR than Fisher, GA, and LoReUn while preserving clean utility, and delivers a 77× wall‑clock speedup over clean‑corpus LoRA retraining, with additional evaluations supporting its effectiveness.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Computation and Language.

arXiv Machine Learning
Jul 15

Inference-Time Machine Unlearning via Gated Activation Redirection

arXiv:2605. 12765v3 Announce Type: replace Abstract: Large Language Models memorize vast amounts of training data, raising concerns regarding privacy, copyright infringement, and safety.

By Vin\'icius Conte Turani, Ot\'avio Parraga, Jo\~ao Vitor Boer Abitante, Kristen K. Arguello, Joana Pasquali, Ramiro N. Barros, Flavio du Pin Calmon, Christian Mattjie, Rodrigo C. Barros, Lucas S. Kupssinsk\"u
arXiv Computation and Language
3d ago

DataFlex: A Unified Framework for Data-Centric Dynamic Training of Large Language Models

arXiv:2603.26164v2 Announce Type: replace-cross Abstract: Data-centric training has emerged as a promising direction for improving large language models (LLMs) by optimizing not only model parameters...

By Hao Liang, Zhengyang Zhao, Mingrui Chen, Meiyi Qiang, Lu Ma, Rongyi Yu, Hengyi Feng, Shixuan Sun, Zimo Meng, Xiaochen Ma, Xuanlin Yang, Qifeng Cai, Ruichuan An, Bohan Zeng, Zhen Hao Wong, Chengyu Shen, Runming He, Zhaoyang Han, Yaowei Zheng, Fangcheng Fu, Conghui He, Bin Cui, Zhiyu Li, Weinan E, Wentao Zhang
arXiv AI
2d ago

Backdoor Purification for LoRA-Tuned LLMs via Null-Space Projection

The paper introduces a method to purify LoRA-tuned large language models (LLMs) against backdoor attacks without relying on trigger knowledge, clean references, or retraining. By extracting high‑fidelity backdoor directions and projecting LoRA updates onto orthogonal null spaces in input and output channels, the approach reduces attack success rates from nearly 100% to under 10%. Experiments demonstrate that this null‑space projection preserves both the base model’s general capabilities and the new downstream skills learned through the adapter across various tasks.

By Jianwei Li, Jung-Eun Kim