Hugging Face Trending Papers

L1 Augmented Attention as an Improved Vector Similarity Metric

Read the original on Hugging Face Trending Papers →

Scaled dot product attention conflates directional alignment and vector magnitude, limiting its effectiveness as a similarity metric in Transformer models. We introduce L1 augmented attention, a simple and computationally parallelizable modification that subtracts a learned, head specific L1 distance between queries and keys from the dot product score.

Summary generated by The Flow from the publisher's feed. The full article lives at Hugging Face Trending Papers.