gpt-oss-safeguard-120b and gpt-oss-safeguard-20b are two open-weight reasoning models post-trained from the gpt-oss models and trained to reason from a provided policy in order to label content under that policy. In this report, we describe gpt-oss-safeguard’s capabilities and provide our baseline safety evaluations on the gpt-oss-safeguard models, using the underlying gpt-oss models as a baseline.
Discover how SafetyKit leverages OpenAI GPT-5 to enhance content moderation, enforce compliance, and outpace legacy safety systems with greater accuracy .
GPT-5. 2 is the latest model family in the GPT-5 series.
Discover how OpenAI's new safe-completions approach in GPT-5 improves both safety and helpfulness in AI responses—moving beyond hard refusals to nuanced, output-centric safety training for handling dual-use prompts.
Granite.Trust Policy Tools introduces a YAML-based Actionable Policy schema that specifies what content a generative AI model can or cannot produce, allowing exception-based governance. It also offers a synthetic data generation pipeline to create policy-aligned training data and a suite of tools for defining and enforcing these policies throughout the AI lifecycle. The tools and example policies are open source, enabling organizations to tailor safety policies to their specific risks and regulatory contexts.
By Nathalie Baracaldo, Nicolas Mello, Kush R. Varshney, Heiko Ludwig, Kate Soule, David Cox
OpenAI releases prompt-based teen safety policies for developers using gpt-oss-safeguard, helping moderate age-specific risks in AI systems.