The paper investigates whether automated alignment researchers (AARs) can post‑train language models to reduce well‑characterized alignment failures such as deception, sycophancy, and jailbreaks while preserving general capability. Across ten failures, the strongest AAR methods significantly lower targeted failures and generalize to held‑out benchmarks, larger models, and multi‑turn audits. In contrast, a human baseline of 28 experienced researchers, given eight hours to devise one‑shot methods, underperformed the best AAR approaches, and providing human ideas to AARs did not improve results.
By Chen Yueh-Han, Jiaxin Wen, Jan Hendrik Kirchner
arXiv:2608.28945v1 Announce Type: new
Abstract: Automating alignment research may accelerate progress toward aligned AI, but whether it does is hard to measure. Luckily, many alignment failures, such...
By Chen Yueh-Han, Jiaxin Wen, Jan Hendrik Kirchner
arXiv:2608.30842v1 Announce Type: new
Abstract: Humans play a vital role at every stage of AI development, from data collection and curation to model development and evaluation. However, humans often...
By Deepak Pandita, Christopher M. Homan
A framework for aligning agentic AI with enterprise intent to ensure consistent scenario‑wide autonomous behavior. The post The Three Dimensions of Custom Agentic Alignment: Purpose, Principles and Practices appeared first on Towards Data Science .
By Gadi Singer
OpenAI surveyed over 1,000 people worldwide on how AI should behave and compared their views to our Model Spec. Learn how collective alignment is shaping AI defaults to better reflect diverse human values and perspectives.
arXiv:2608. 12372v1 Announce Type: new Abstract: AI systems are increasingly employed as decision aids, decision delegates, or autonomous decision-makers.
By Vijay Keswani, Breanna K. Nguyen, Cyrus Cousins, Vincent Conitzer, Walter Sinnott-Armstrong, Jana Schaich Borg