arXiv AI By Haoran Zhang, Feng Zhou

Flexformer: Flexible Linear Transformer with Learnable Attention Kernel

Read the original on arXiv AI →

arXiv:2606. 27748v1 Announce Type: cross Abstract: Transformer models rely on attention mechanism to capture long-range dependencies but suffer from quadratic complexity, limiting their scalability to long sequences.

Summary generated by The Flow from the publisher's feed. The full article lives at arXiv AI.