Learning What Matters: Supervising Global Context Pruning with Causal Evidence Sets
Read the original on arXiv Machine Learning →arXiv:2607. 21692v2 Announce Type: replace Abstract: Sparse attention prunes a long context to the blocks a model needs, and the usual selector is distilled from a dense teacher's attention.
Summary generated by The Flow from the publisher's feed. The full article lives at arXiv Machine Learning.