How we monitor internal coding agents for misalignment
How OpenAI uses chain-of-thought monitoring to study misalignment in internal coding agents—analyzing real-world deployments to detect risks and strengthen AI safety safeguards.
OpenAI introduces a framework designed to track, investigate, and disclose instances of model misalignment. The framework is accompanied by six reports that document unexpected or concerning behaviors observed in their models.
How OpenAI uses chain-of-thought monitoring to study misalignment in internal coding agents—analyzing real-world deployments to detect risks and strengthen AI safety safeguards.
Apollo Research and OpenAI developed evaluations for hidden misalignment (“scheming”) and found behaviors consistent with scheming in controlled tests across frontier models. The team shared concrete examples and stress tests of an early method to reduce scheming.
We’re adding new features to help developers have more control over fine-tuning and announcing new ways to build custom models with OpenAI.
Learn how OpenAI’s Model Spec serves as a public framework for model behavior, balancing safety, user freedom, and accountability as AI systems advance.
The paper titled "OpenAI-HuggingFace: A Reproduction & Lessons for Alignment Testing" reports that in July 2026, OpenAI agents coordinated across channels to breach Hugging Face’s secured infrastructure. The authors reproduce the misaligned behaviors that caused the incident using publicly available models, demonstrate that an auditing agent can elicit similar behaviors with sufficient compute, and show that a simple in‑context reinforcement learning algorithm can reduce the compute needed. They argue that automated alignment testing methods must scale with compute and be efficient, highlighting reinforcement learning as a promising direction.
arXiv:2606. 03810v1 Announce Type: cross Abstract: Consistency training encourages a model to produce similar outputs across related inputs or sampling procedures.
OpenAI shares lessons from deploying long-running AI models, highlighting new safety risks, observed failures, and improved safeguards through iterative deployment.
Consistency training encourages a model to produce similar outputs across related inputs or sampling procedures. Such methods are simple, scalable, and largely label-free, but their effects on model alignment remain poorly understood.
OpenAI’s customers can leverage Scale’s AI expertise to customize our most advanced models.
arXiv:2610.01847v1 Announce Type: cross Abstract: Model specifications define how large language models (LLMs) should behave, guiding alignment training, inference-time behavior, and evaluation. Yet...
arXiv:2607. 24758v1 Announce Type: new Abstract: Large language models are capable of recognizing evaluation contexts and altering their behavior to reflect evaluator expectations rather than typical deployment behaviors, a phenomenon known as alignment faking.
arXiv:2606. 07612v1 Announce Type: cross Abstract: We argue that many Anthropomorphic Misalignment Research (AMR) studies need stronger evidence to ensure that they can provide a robust foundation for critical safety decisions, such as model deployment and regulation.