arXiv AI

A Coulomb Particle Model for Learning Kernel Attention in Transformers

arXiv:2607. 23869v1 Announce Type: cross Abstract: Randomized features provide a scalable approximation to kernel machines, but their performance depends strongly on the choice of feature distribution.

Hugging Face Trending Papers
Jul 20

L1 Augmented Attention as an Improved Vector Similarity Metric

Scaled dot product attention conflates directional alignment and vector magnitude, limiting its effectiveness as a similarity metric in Transformer models. We introduce L1 augmented attention, a simple and computationally parallelizable modification that subtracts a learned, head specific L1 distance between queries and keys from the dot product score.