How Do Transformers Learn to Represent Symmetries?
Read the original on arXiv Machine Learning →The paper investigates how a vanilla Transformer learns symmetries from finite data augmentation on point cloud datasets. It finds an ordering of learnability: non-angle-preserving symmetries are easiest, followed by angle-preserving symmetries, and finally base angle-preserving subgroups such as translation, rotation, and scale. The study also examines the Transformer's extrapolation behavior, performs a structural analysis of trained models to uncover interpretable mechanisms for invariance, and extends these findings to equivariant functions, suggesting that the identified mechanisms can serve as building blocks for learned equivariance.
Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.