arXiv Machine Learning

Model Forensics: Investigating Whether Concerning Behavior Reflects Misalignment

arXiv:2606. 26071v1 Announce Type: new Abstract: A central goal of safety research is determining whether a model is misaligned.

arXiv AI
6d ago

Backdoor Decontamination Dynamics in LLM Agents

arXiv:2608. 11295v1 Announce Type: cross Abstract: Open-weight LLM agents are vulnerable to backdoors installed during fine-tuning, which may be undetectable if the trigger conditions are never met during testing.

By Gabriel Huang, Abhay Puri, L\'eo Boisvert, Alexandre Drouin, Perouz Taslakian, Spandana Gella, Christopher Pal