arXiv AI By Omar Mahmoud, Aly M. Kassem, Thommen George Karimpanal, Buddhika Laknath Semage, Negar Rostamzadeh, Golnoosh Farnadi, Santu Rana

Shared Latent Structures Enable Unified Backdoor Detection and Mitigation in LLMs

Read the original on arXiv AI →

arXiv:2606. 07963v1 Announce Type: new Abstract: Backdoor attacks in large language models (LLMs) are often treated as isolated trigger-response failures, motivating defenses tailored to specific triggers or behaviors.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Machine Learning
Jul 30

ToxScreen: Detecting Whether an LLM Has Been Poisoned

arXiv:2607. 26849v1 Announce Type: cross Abstract: As large language models (LLMs) are deployed in high-stakes domains, adversaries may poison training data to implant backdoors: hidden triggers that covertly manipulate model behavior at inference time.

By Anthony Hughes, Nicole Xing, Collin Francel, Andy Kim, Andrew Draganov
arXiv Computation and Language
Sep 25

LLM Forensics: Where Do Backdoors Hide? Localizing and Controlling Trigger Mechanisms with Sparse Autoencoders

The paper investigates how trigger-based backdoors function in large language models by using sparse autoencoders (SAEs) to identify feature directions across layers and transformer components. In a controlled experiment, 1B and 8B models were made to continue English prompts in French or German when presented with fixed trigger sequences. The study finds that different SAE feature directions correspond to trigger detection, residual-stream propagation, and language tracking, with residual-stream features being most effective for controlling the backdoor behavior.

By Wissam Antoun, Francis Kulumba, Th\'eo Lasnier, Beno\^it Sagot, Djam\'e Seddah