OpenAI Blog

Introducing Activation Atlases

We’ve created activation atlases (in collaboration with Google researchers), a new technique for visualizing what interactions between neurons can represent. As AI systems are deployed in increasingly sensitive contexts, having a better understanding of their internal decision-making processes will let us identify weaknesses and investigate failures.

OpenAI Blog
Apr 14, 2020

OpenAI Microscope

We’re introducing OpenAI Microscope, a collection of visualizations of every significant layer and neuron of eight vision “model organisms” which are often studied in interpretability. Microscope makes it easier to analyze the features that form inside these neural networks, and we hope it will help the research community as we move towards understanding these complicated systems.

Hugging Face Trending Papers
Aug 19

Graphical Design of Interpretable Architectures

The paper introduces a new graphical notation, adapted from Penrose tensor notation, to design and represent interpretable AI architectures. Unlike symbolic equations or probabilistic models, this notation provides a global view of an architecture while directly mapping to PyTorch einsum code. The authors demonstrate its use on several interpretable models and on the Steerling-8B language model, revealing structural insights and enabling concise code generation.

arXiv AI
Aug 20

Graphical Design of Interpretable Architectures

The paper introduces a graphical notation, adapted from Penrose tensor notation, to design and represent interpretable AI architectures. This notation provides a global view of an architecture and maps directly onto PyTorch einsum code, enabling clear depiction of tensor manipulations. The authors apply the notation to several interpretable models—concept bottlenecks, sparse probes, prototype networks, neural additive models, and mixtures of linear models—and use it to diagram the key components of the Steerling-8B language model, revealing its residual structure and allowing a concise 33‑line PyTorch implementation.

By Pietro Barbiero
arXiv Machine Learning
Sep 4

When Vision Meets Graphs: A Survey on Graph Reasoning and Learning

The survey titled "When Vision Meets Graphs: A Survey on Graph Reasoning and Learning" reviews how visual depictions of graphs can be used as inputs for graph reasoning and learning. It highlights that while Graph Neural Networks dominate graph machine learning, most pipelines ignore the visual form of graphs, despite scientists routinely interpreting graphs visually. The paper organizes existing work into three threads—vision for graph reasoning, vision for graph learning, and scientific graphs—aiming to clarify current capabilities and chart a path toward foundation models that perceive and reason about graphs like scientists do.

By Xinjian Zhao, Wei Pang, Zhixuan Yu, Xiangru Jian, Xiaozhuang Song, Yaoyao Xu, Zhongkai Xue, Dingshuo Chen, Shu Wu, Philip Torr, Tianshu Yu
arXiv Machine Learning
Jul 7

On a Geometry of Interbrain Networks

arXiv:2509. 10650v4 Announce Type: replace-cross Abstract: Effective analysis in neuroscience benefits significantly from robust conceptual frameworks.

By Nicol\'as Hinrichs, Noah Guzm\'an, Melanie Weber