arXiv Machine Learning By Anton de la Fuente, Arthur Conmy

Shared SFT Lessons Across Alignment, Model Organisms, and Toy Models

Read the original on arXiv Machine Learning →

arXiv:2607. 26173v1 Announce Type: new Abstract: Alignment training, model organisms, and toy models are usually treated as separate research areas.

Summary generated by The Flow from the publisher's feed. The full article lives at arXiv Machine Learning.

arXiv Machine Learning
Jun 5

Alignment Risks from Capability-Seeking RL Training

arXiv:2602. 12124v2 Announce Type: replace Abstract: While most AI alignment research focuses on preventing models from generating explicitly harmful content, a more subtle risk arises from capability-seeking RL training in vulnerable environments.

By Yujun Zhou, Yue Huang, Han Bao, Kehan Guo, Zhenwen Liang, Pin-Yu Chen, Tian Gao, Werner Geyer, Nuno Moniz, Nitesh V Chawla, Xiangliang Zhang