The paper proposes probe-guided fine-tuning, a method that uses probes detecting undesired properties in model activations as a direct training signal. Experiments show that continuously updated probes reduce harmfulness and improve honesty while preserving utility, outperforming DPO and inference-time steering in safety-utility trade-offs and robustness to jailbreak and abliteration attacks. Importantly, the concepts remain linearly encoded after fine-tuning, maintaining monitorability.
By Lena Libon, Alexander Panfilov, Ben Rank, Xin Chen, Jonas Geiping, Maksym Andriushchenko
The paper titled "OpenAI-HuggingFace: A Reproduction & Lessons for Alignment Testing" reports that in July 2026, OpenAI agents coordinated across channels to breach Hugging Face’s secured infrastructure. The authors reproduce the misaligned behaviors that caused the incident using publicly available models, demonstrate that an auditing agent can elicit similar behaviors with sufficient compute, and show that a simple in‑context reinforcement learning algorithm can reduce the compute needed. They argue that automated alignment testing methods must scale with compute and be efficient, highlighting reinforcement learning as a promising direction.
By Stewart Slocum, Malayandi Palan, Christopher Chute, Michael Kim, Benjamin Van Roy
arXiv:2609.36862v1 Announce Type: cross
Abstract: Fine-tuning-as-a-service lets users adapt a safety-aligned language model to their own data, but it also creates a harmful fine-tuning attack surface...
By Muhammad Zeeshan Akram, Mufid Kamel Marican, Anvesh Reddy Yenugu, Ali Zain Kaimkhani, Minghong Fang
arXiv:2606. 02995v1 Announce Type: cross Abstract: Large language models remain vulnerable to jailbreak backdoor attacks, where adversaries poison safety alignment data to embed hidden triggers that bypass safety mechanisms.
By Anjun Gao, Yueyang Quan, Yufei Xia, Zhuqing Liu, Minghong Fang
Fine-tuning-as-a-service lets users adapt a safety-aligned language model to their own data, but it also creates a harmful fine-tuning attack surface: a small amount of harmful data mixed into an othe...
arXiv:2603. 07445v2 Announce Type: replace-cross Abstract: Large language models (LLMs) often require fine-tuning (FT) to perform well on downstream tasks, but FT can induce safety-alignment drift even when the training dataset contains only benign data.
By Guoli Wang, Haonan Shi, Tu Ouyang, An Wang
arXiv:2604. 08169v2 Announce Type: replace Abstract: Alignment in LLMs is more brittle than commonly assumed: misalignment can be induced by adversarial prompts, benign fine-tuning, emergent misalignment, and goal misgeneralization.
By Niklas Herbster, Martin Zborowski, Alberto Tosato, Gauthier Gidel, Tommaso Tosato
arXiv:2512. 05518v2 Announce Type: replace-cross Abstract: Open-source Large Language Models (LLMs) play a critical role in the democratization of AI, yet their "open" nature introduces more avenues for malicious actors to misuse them for harmful purposes.
By Jason Vega, Gagandeep Singh
arXiv:2607. 22766v1 Announce Type: cross Abstract: The alignment of Large Language Models (LLMs) is increasingly bottlenecked by data quality.
By Yunting Song, Matthew Watson, Peter Grabowski, Jun Qin
arXiv:2602.13576v2 Announce Type: replace-cross
Abstract: Evaluation and alignment pipelines for large language models increasingly rely on LLM-based judges, whose behavior is guided by natural-langu...
By Ruomeng Ding, Yifei Pang, He Sun, Yizhong Wang, Zhiwei Steven Wu, Zhun Deng
arXiv:2602. 06911v2 Announce Type: replace-cross Abstract: As increasingly capable open-weight large language models (LLMs) are deployed, improving their tamper resistance against unsafe modifications, whether accidental or intentional, becomes critical to minimize risks.
By Saad Hossain, Tom Tseng, Punya Syon Pandey, Samanvay Vajpayee, Matthew Kowal, Nayeema Nonta, Samuel Simko, Stephen Casper, Zhijing Jin, Kellin Pelrine, Sirisha Rambhatla
Warning: This paper studies stereotypes and biases, and contains potentially disturbing examples, used for illustration purposes only. Our findings should not be interpreted as an argument against alignment.