arXiv Statistics ML By Demi\'an Fraiman

On the Expressive Power of Transformers for Contextual Relations

Read the original on arXiv Statistics ML →

The paper investigates the theoretical expressive power of Transformers in modeling contextual relations. By framing a text as a distribution of representations and attention as a probabilistic relation, it connects attention normalization to optimal transport: softmax yields conditional relations, while Sinkhorn yields joint relations with fixed marginals. The authors prove universal approximation results, showing that Transformers with Sinkhorn normalization can represent any joint probability relation, whereas standard softmax Transformers can represent any conditional probability relation.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Statistics ML.

arXiv Machine Learning
Sep 10

Conditioned Initialization for Attention

arXiv:2609.07086v1 Announce Type: new Abstract: Transformers are a dominant architecture in modern machine learning, powering applications across vision, language, and beyond. At the core of their su...

By Hemanth Saratchandran, Simon Lucey