arXiv AI By David N. Olivieri, Antonio F. P\'erez Rodr\'iguez

Transformer Field Theory: A Response-Theoretic Approach to Mechanistic Interpretability

Read the original on arXiv AI →

arXiv:2605. 25225v2 Announce Type: replace-cross Abstract: Mechanistic interpretability often studies Transformer behavior by intervening on internal activations through activation patching, causal tracing, path patching, and steering directions.

Summary generated by The Flow from the publisher's feed. The full article lives at arXiv AI.