Hierarchical Latent Structures in Data Generation Process Unify Mechanistic Phenomena across Scale
Read the original on arXiv Machine Learning →arXiv:2603. 06592v2 Announce Type: replace-cross Abstract: Contemporary studies in mechanistic interpretability have uncovered many puzzling phenomena in the neural information processing of Transformer-based language models, such as induction heads, function vectors, and the Hydra effect.
Summary generated by The Flow from the publisher's feed. The full article lives at arXiv Machine Learning.