arXiv Machine Learning By Dhruman Gupta, Aritra Das, Debayan Gupta

The Query Knows What to Forget: A Second Erase Direction for Linear Attention

Read the original on arXiv Machine Learning →

arXiv:2608. 13668v1 Announce Type: new Abstract: Linear attention keeps a state of fixed size.

Summary generated by The Flow from the publisher's feed. The full article lives at arXiv Machine Learning.

arXiv Machine Learning
Jun 16

KVEraser: Learning to Steer KV Cache for Efficient Localized Context Erasing

arXiv:2606. 17034v1 Announce Type: cross Abstract: Post-hoc context erasing over the KV cache is challenging because a local edit has a global consequence: once a span has been processed, its influence propagates into the cached states of all subsequent tokens.

By Mufei Li, Shikun Liu, Dongqi Fu, Haoyu Wang, Yinglong Xia, Hong Li, Hong Yan, Pan Li