arXiv:2606. 26523v1 Announce Type: new Abstract: We develop a framework for interpreting AI systems as agents, drawing on the philosophical tradition of radical interpretation and the tools of mechanistic interpretability.
By Daniel A. Herrmann, Benjamin A. Levinstein
The paper argues that as Large Language Models transition from chatbots to agentic systems, the current post-hoc interpretability paradigm is insufficient for safe deployment because it cannot audit or intervene before an output is produced. It proposes a shift to generative interpretability, where a model’s inference process inherently exposes semantically meaningful checkpoints that are human-understandable and can be causally intervened upon. The authors illustrate the advantages of this approach and introduce Neuro‑Symbolic Models as a concrete implementation.
By Xiaocong Yang
arXiv:2606. 19857v1 Announce Type: cross Abstract: Large language models (LLMs) are commonly prompted and interfaced with human-readable natural language, even when the intended reader is another model.
By Jiayi Zhu, Haoxuan Peng, Junxi Wang, Liang Ke, Chen Zhang, Linfeng Zhang
arXiv:2608. 19579v1 Announce Type: new Abstract: Large Language Models (LLMs) are increasingly deployed in high-stakes applications, yet their tendency to generate toxic, harmful, or policy-violating content poses significant risks.
By Mohamed Akrout, Olivera Kotevska, Dan Wilson
arXiv:2606. 23991v1 Announce Type: new Abstract: What is an agent?
By Eric Xing, Mingkai Deng, Jinyu Hou
The paper extends a dynamical systems approach to classify unsafe outputs from large language models (LLMs) by projecting prompts and responses into high‑dimensional embeddings and fitting separate Koopman-based predictive models for safe and unsafe regimes. A differential residual score compares prediction errors from these models to classify new outputs. Experiments on three safety benchmarks show that including prompt embeddings improves detection of interaction‑dependent violations, especially with causal decoders like Llama‑3, while response‑only violations benefit more from dense semantic embeddings.