DeepMind Blog

Gemma Scope 2: helping the AI safety community deepen understanding of complex language model behavior

Read the original on DeepMind Blog →

Open interpretability tools for language models are now available across the entire Gemma 3 family with the release of Gemma Scope 2.

Summary generated by The Flow from the publisher's feed. The full article lives at DeepMind Blog.

arXiv AI
Aug 11

SHE: Trajectory-driven Safety Harness Evolution for LLM Agents

arXiv:2608. 09885v1 Announce Type: new Abstract: The safety of large language model (LLM) agents depends not only on model weights but also on the agent harness that manages context, memory, tools, permissions, and runtime control.

By Wanying Qu, Qinghua Mao, Yu Li, Jiyao Liu, Xin Zhang, Dadi Guo, Yanxu Zhu, Qingyu Liu, Leitao Yuan, Xi Lin, Shanfeng Zhu, Yanwei Fu, Jing Shao, Xia Hu, Dongrui Liu