arXiv Machine Learning By Dimitris Dimakopoulos, Shay B. Cohen, Ioannis Konstas

MoRFI: Monotonic Sparse Autoencoder Feature Identification

Read the original on arXiv Machine Learning →

arXiv:2604. 26866v2 Announce Type: replace-cross Abstract: Large language models (LLMs) acquire most of their factual knowledge during the pre-training stage, through next token prediction.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv Machine Learning
Sep 16

Test-Time Unlearning via Sparse Autoencoder

arXiv:2609.16229v1 Announce Type: new Abstract: Machine unlearning aims to remove specific knowledge from a trained large language model (LLM) without retraining from scratch. Existing methods modify...

By Pingzhi Li, Jinhao Duan, Vaishnav Tadiparthi, Nakul Agarwal, Kwonjoon Lee, Ehsan Moradi Pari, Hossein Nourkhiz Mahjoub, Sijia Liu, Tianlong Chen
arXiv AI
Sep 2

Why Fine-Tuning Encourages Hallucinations and How to Fix It

The paper investigates why supervised fine‑tuning (SFT) of large language models leads to increased hallucinations of factual information. It proposes a self‑distillation SFT approach that regularizes output‑distribution drift to preserve pre‑training knowledge, and shows that freezing parameter groups can reduce hallucinations when new knowledge is unnecessary. Experiments attribute the main cause to interference among overlapping semantic representations, which self‑distillation mitigates, and an associative‑memory model explains the forgetting dynamics.

By Guy Kaplan, Zorik Gekhman, Zhen Zhu, Lotem Rozner, Yuval Reif, Swabha Swayamdipta, Derek Hoiem, Roy Schwartz
arXiv AI
Jun 9

Shared Latent Structures Enable Unified Backdoor Detection and Mitigation in LLMs

arXiv:2606. 07963v1 Announce Type: new Abstract: Backdoor attacks in large language models (LLMs) are often treated as isolated trigger-response failures, motivating defenses tailored to specific triggers or behaviors.

By Omar Mahmoud, Aly M. Kassem, Thommen George Karimpanal, Buddhika Laknath Semage, Negar Rostamzadeh, Golnoosh Farnadi, Santu Rana