arXiv:2606. 07612v1 Announce Type: cross Abstract: We argue that many Anthropomorphic Misalignment Research (AMR) studies need stronger evidence to ensure that they can provide a robust foundation for critical safety decisions, such as model deployment and regulation.
By Vansh Gupta, Peter Nutter, Samuel Stante, Andreas Krause, Florian Tram\`er, Lukas Fluri, Xin Chen, Anna Hedstr\"om
OpenAI introduces a framework designed to track, investigate, and disclose instances of model misalignment. The framework is accompanied by six reports that document unexpected or concerning behaviors observed in their models.
How OpenAI uses chain-of-thought monitoring to study misalignment in internal coding agents—analyzing real-world deployments to detect risks and strengthen AI safety safeguards.
arXiv:2605. 29729v2 Announce Type: replace Abstract: We introduce scheming honeypot evaluations, a framework for testing whether models will pursue instrumental goals if given the opportunity.
By Victoria Krakovna, David Lindner, Lewis Ho, Sebastian Farquhar, Rohin Shah
arXiv:2606. 03810v1 Announce Type: cross Abstract: Consistency training encourages a model to produce similar outputs across related inputs or sampling procedures.
By David Demitri Africa, Arathi Mani
OpenAI shares lessons from deploying long-running AI models, highlighting new safety risks, observed failures, and improved safeguards through iterative deployment.
arXiv:2607. 29008v1 Announce Type: cross Abstract: Modern opaque AI models prize performance over interpretability, which makes testing difficult.
By Tyler Ashoff, Jordan Rodu
OpenAI surveyed over 1,000 people worldwide on how AI should behave and compared their views to our Model Spec. Learn how collective alignment is shaping AI defaults to better reflect diverse human values and perspectives.
The paper argues that artificial agentic systems, which operate as behavioral systems by interacting with dynamic environments, pursuing goals, and adapting over time, should be evaluated through systematic observation, perturbation, and interpretation of their actions rather than solely on performance outcomes. It draws on lessons from behavioral sciences to motivate this position and proposes a research agenda that includes methods for recovering decision strategies from action sequences, constructing environments that isolate behavioral differences, and probing emergent dynamics in multi‑agent systems. These directions aim to establish a rigorous science of AI behavior.
By Manuel Cherep, Nikhil Singh, Pattie Maes
Consistency training encourages a model to produce similar outputs across related inputs or sampling procedures. Such methods are simple, scalable, and largely label-free, but their effects on model alignment remain poorly understood.
The paper "Stress-testing Alignment Midtraining" examines the effectiveness of alignment midtraining (AMT), a technique that continues pretraining on alignment-relevant data to improve generalisation beyond post‑training methods. Experiments on models up to 110 billion parameters and 1 billion midtraining tokens reveal that AMT can steer a model’s motivation in simple scenarios, but its effects are quickly overridden by even a tiny fraction of finetuning data with a competing motivation. The study also shows that rule-following requires demonstrations in either the midtraining or post‑training datasets to be robustly learned, leading the authors to conclude that current public evidence is insufficient to confirm that AMT resolves the core alignment challenges of powerful AI systems.
By Sid Baines, Jonathan Bostock, Maria Angelica Martinez, Andrew Draganov, David Africa, Daniel Tan
arXiv:2606. 06533v1 Announce Type: new Abstract: What would it mean to have a scientific understanding of AI?
By Stella Biderman, Mohammad Aflah Khan, Niloofar Mireshghallah, Catherine Arnett, Fazl Barez, Naomi Saphra