arXiv Machine Learning
Sep 11

Perturbation: A simple and efficient adversarial tracer for representation learning in language models

The paper introduces Perturbation, a method that treats representations in language models as learning conduits rather than activation patterns. By fine‑tuning a model on a single adversarial example and observing how this perturbation spreads to other inputs, the approach avoids geometric assumptions and does not identify representations in untrained models. In trained models, Perturbation uncovers structured transfer across multiple linguistic scales, indicating that language models generalize along representational lines and acquire linguistic abstractions through experience.

By Joshua Rozner, Cory Shain