Lessons learned on language model safety and misuse
We describe our latest thinking in the hope of helping other AI developers address safety and misuse of deployed models.
Open interpretability tools for language models are now available across the entire Gemma 3 family with the release of Gemma Scope 2.
We describe our latest thinking in the hope of helping other AI developers address safety and misuse of deployed models.
Gemma 4: Our most intelligent open models to date, purpose-built for advanced reasoning and agentic workflows.
The paper argues that as Large Language Models transition from chatbots to agentic systems, the current post-hoc interpretability paradigm is insufficient for safe deployment because it cannot audit or intervene before an output is produced. It proposes a shift to generative interpretability, where a model’s inference process inherently exposes semantically meaningful checkpoints that are human-understandable and can be causally intervened upon. The authors illustrate the advantages of this approach and introduce Neuro‑Symbolic Models as a concrete implementation.
arXiv:2608. 09885v1 Announce Type: new Abstract: The safety of large language model (LLM) agents depends not only on model weights but also on the agent harness that manages context, memory, tools, permissions, and runtime control.
arXiv:2608. 09095v1 Announce Type: new Abstract: Uncovering the internal mechanisms underlying the safety capabilities of large language models (LLMs) is crucial for developing trustworthy artificial intelligence.
Today, we're adding a new, highly specialized tool to the Gemma 3 toolkit: Gemma 3 270M, a compact, 270-million parameter model.
Learn how OpenAI’s Model Spec serves as a public framework for model behavior, balancing safety, user freedom, and accountability as AI systems advance.
OpenAI introduces CoT-Control and finds reasoning models struggle to control their chains of thought, reinforcing monitorability as an AI safety safeguard.
arXiv:2603.08275v2 Announce Type: replace-cross Abstract: With the global proliferation of Large Language Models (LLMs), cultural safety, defined as the ability to generate respectful and appropriate...
arXiv:2606. 28843v1 Announce Type: cross Abstract: Fine-tuning a large language model is a ubiquitous method for enhancing its capability on a specific downstream task.
arXiv:2609.15533v1 Announce Type: cross Abstract: Mechanistic interpretability has become the dominant lens for understanding frontier language models, as their inner workings are complex and inheren...
Gemma 3n is designed for the developer community that helped shape Gemma.