Our approach to AI safety
Ensuring that AI systems are built, deployed, and used safely is critical to our mission.
How we think about safety for users experiencing mental or emotional distress, the limits of today’s systems, and the work underway to refine them.
Ensuring that AI systems are built, deployed, and used safely is critical to our mission.
An update on our safety & security practices
OpenAI collaborated with 170+ mental health experts to improve ChatGPT’s ability to recognize distress, respond empathetically, and guide users toward real-world support—reducing unsafe responses by up to 80%. Learn how we’re making ChatGPT safer and more supportive in sensitive moments.
arXiv:2608. 07902v1 Announce Type: cross Abstract: Youth increasingly turn to AI chatbots for social and emotional support, raising concerns about how these systems respond, especially in high-stakes situations.
arXiv:2607. 22692v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used for emotional support despite lacking mechanisms to safely govern evolving mental health risk.
We’re forming a new industry body to promote the safe and responsible development of frontier AI systems: advancing AI safety research, identifying best practices and standards, and facilitating information sharing among policymakers and industry.
arXiv:2607. 14353v1 Announce Type: cross Abstract: As automated decision-making and data-driven technologies pervade society and are used to manage consequential outcomes, understanding the technology's capabilities, limitations, and attendant risks in context requires analysis of full sociotechnical systems.
Learn how new ChatGPT safety updates improve context awareness in sensitive conversations, helping detect risk over time and respond more safely.
Anian is a safety‑gated multimodal AI backend designed for perinatal mental‑health support and mindfulness‑intervention routing. It maps user input into a four‑layer hierarchical state representation—emotion, psychosocial constructs, safety risk, and intervention routes—then fuses local and external risk signals to decide whether to generate AI responses or provide fixed safety content. Prototype evaluation on large public corpora showed high classification performance and perfect high‑risk recall in a controlled stress test, though clinical validity remains unestablished.
arXiv:2606. 23884v1 Announce Type: cross Abstract: General-purpose large language models (LLMs) are increasingly used for mental health-related conversations, yet safety safeguards remain inadequate and inconsistent across clinical conditions.