arXiv AI By Lena Libon, Alexander Panfilov, Ben Rank, Xin Chen, Jonas Geiping, Maksym Andriushchenko

Alignment via Training Against Probes Without Losing Monitorability

Read the original on arXiv AI →

The paper proposes probe-guided fine-tuning, a method that uses probes detecting undesired properties in model activations as a direct training signal. Experiments show that continuously updated probes reduce harmfulness and improve honesty while preserving utility, outperforming DPO and inference-time steering in safety-utility trade-offs and robustness to jailbreak and abliteration attacks. Importantly, the concepts remain linearly encoded after fine-tuning, maintaining monitorability.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Machine Learning
Jun 5

Alignment Risks from Capability-Seeking RL Training

arXiv:2602. 12124v2 Announce Type: replace Abstract: While most AI alignment research focuses on preventing models from generating explicitly harmful content, a more subtle risk arises from capability-seeking RL training in vulnerable environments.

By Yujun Zhou, Yue Huang, Han Bao, Kehan Guo, Zhenwen Liang, Pin-Yu Chen, Tian Gao, Werner Geyer, Nuno Moniz, Nitesh V Chawla, Xiangliang Zhang
arXiv Computation and Language
Sep 4

Beyond Shallow Alignment: How Post-Training Methods Determine Refusal Circuits And Steering Robustness

The paper investigates how different post‑training methods—supervised fine‑tuning, reasoning‑augmented fine‑tuning, and preference optimization (ORPO)—affect the internal computation of refusal behavior in language models. Experiments on Llama‑3.1‑8B, Gemma‑2‑9B, and Qwen3‑8B show that reasoning‑augmented training consistently creates a distinct refusal computation across models, while the architecture influences the internal structure and steerability of refusal. None of the studied methods simultaneously achieve a distributed refusal mechanism, preserve general capability, and allow easy corrective edits, indicating that current post‑training approaches are not a fully reliable defense for safety-critical applications.

By Hoang Cuong Nguyen, Mark Dras, Usman Naseem