arXiv AI

Transcoders for Investigating Deception in Language Models

arXiv:2607. 14791v1 Announce Type: new Abstract: Transcoders have recently emerged as a promising approach for mechanistic interpretability (MI), enabling circuit-level analysis of model behaviour.

arXiv Machine Learning
Jun 4

Covert Influence Between Language Models

arXiv:2606. 04071v1 Announce Type: cross Abstract: As language models increasingly consume one another's outputs, covert influence -- a phenomenon where a sender's payload (the behavioral disposition it is conditioned to propagate) transfers to a receiver through carriers undetectable by humans -- becomes a growing risk.

By Avidan Shah, Jay Chooi, Jinghua Ou, Shi Feng
arXiv AI
Jun 9

Shared Latent Structures Enable Unified Backdoor Detection and Mitigation in LLMs

arXiv:2606. 07963v1 Announce Type: new Abstract: Backdoor attacks in large language models (LLMs) are often treated as isolated trigger-response failures, motivating defenses tailored to specific triggers or behaviors.

By Omar Mahmoud, Aly M. Kassem, Thommen George Karimpanal, Buddhika Laknath Semage, Negar Rostamzadeh, Golnoosh Farnadi, Santu Rana