Frontier risk and preparedness
To support the safety of highly-capable AI systems, we are developing our approach to catastrophic risk preparedness, including building a Preparedness team and launching a challenge.
This report outlines the safety work carried out prior to releasing deep research including external red teaming, frontier risk evaluations according to our Preparedness Framework, and an overview of the mitigations we built in to address key risk areas.
To support the safety of highly-capable AI systems, we are developing our approach to catastrophic risk preparedness, including building a Preparedness team and launching a challenge.
This report outlines the safety work carried out prior to releasing OpenAI o1 and o1-mini, including external red teaming and frontier risk evaluations according to our Preparedness Framework.
We’re strengthening the Frontier Safety Framework (FSF) to help identify and mitigate severe risks from advanced AI models.
arXiv:2606. 12429v1 Announce Type: cross Abstract: Muse Spark is the latest large language model developed by Meta.
Sharing our updated framework for measuring and protecting against severe harm from frontier AI capabilities.
The article discusses how automated red‑teaming can uncover more vulnerabilities at lower cost than human red‑teaming on AI safety benchmarks, yet this comparison conflates measurement with conclusion. It argues that benchmarks only assess harms within a predefined set, leaving a "threat‑model coverage gap" that can hide new risks, as seen in non‑English prompts. The authors suggest that evaluators from deployment contexts distinct from developers are needed to close this gap.
Google DeepMind and partners announce a $10M funding call for multi-agent safety research.
A pilot program to support independent safety and alignment research and develop the next generation of talent
The OpenAI Blog article titled "Path to Astra: critical capabilities and frontier safeguards" announces that Astra is the first OpenAI model to meet the Critical cybersecurity capability threshold under the Preparedness Framework. It highlights that Astra incorporates stronger safeguards for its release, ensuring higher security standards. The piece underscores the model’s compliance with stringent cybersecurity criteria.
arXiv:2606. 04035v1 Announce Type: cross Abstract: We present a systematic study of domain-dependent safety behavior in open-weight LLMs: 7 standardized experiments across 7 ethical domains, testing 5 models (12B--70B) in 4,200 interactions with dual-judge validation.
Google DeepMind and UK AI Security Institute (AISI) strengthen collaboration on critical AI safety and security research