arXiv AI By David N. Olivieri, Antonio F. P\'erez Rodr\'iguez

Transformer Field Theory: A Response-Theoretic Approach to Mechanistic Interpretability

Read the original on arXiv AI →

arXiv:2605. 25225v2 Announce Type: replace-cross Abstract: Mechanistic interpretability often studies Transformer behavior by intervening on internal activations through activation patching, causal tracing, path patching, and steering directions.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.