arXiv Machine Learning By Vedant Puri, Yongjie Jessica Zhang, Levent Burak Kara

FLARE++: Low-rank attention with dynamic attention routing

Read the original on arXiv Machine Learning →

arXiv:2608. 11519v1 Announce Type: new Abstract: Full self-attention is a strong token mixer for PDE surrogates on irregular domains, but its quadratic cost limits its use on high-resolution problems.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv Machine Learning
Sep 21

Elastic Threshold Attention: Learned Contextual Sparsity for Long-Context Decoding

Elastic Threshold Attention (ETA) is a trainable attention mechanism that dynamically predicts contextual thresholds from query representations, enabling selective pruning of KV cache tokens during long‑context decoding. By multiplicatively suppressing sub‑threshold logits during training, ETA avoids representation collapse and eliminates localized attention sinks, allowing a 1.45B model to match dense attention performance at roughly 85% training sparsity and 38% active decode density. At inference, a custom Triton kernel achieves up to 2.5× faster decoding on sequences up to 512K tokens, and an offline calibration step can further reduce compute by 27% by freezing per‑head thresholds.

By Themistoklis Haris, Henry Li, Maryam Karimzadehgan
arXiv AI
Sep 16

QueryFormer: Winning Solution for KDD Cup 2026 Tencent UniRec Challenge

QueryFormer is a unified architecture designed for post‑click conversion rate prediction, addressing both feature interactions and sequential user behaviors. It introduces a stackable field–sequence block that generates query tokens via cross‑attention and packs sequence queries into shared‑parameter attention, improving efficiency and accuracy. The model won first place in the KDD Cup 2026 Tencent UniRec Challenge Industrial Track with an AUC of 0.83254, and scaling studies show that increasing view width slightly boosts validation AUC while maintaining low latency.

By Yuanzhe Zhou, Zhaoyang Zeng
arXiv Machine Learning
Sep 3

CRISP: Cliff-awaRe Input-adaptive Sparse Prefilling with Structural-Mass-Motivated Routing

CRISP (Cliff-awaRe Input-adaptive Sparse Prefilling) is a new method for long-context LLM inference that replaces costly quadratic attention prefilling with a dynamic, input-adaptive sparse routing scheme. It introduces a structural proxy, C_struct, to directly read routing decisions from the proxy attention map, eliminating the need for pooled matrix multiplication and KL divergence. Additionally, CRISP addresses the post-softmax mass cliff by using a sink-aware threshold based on the noise floor, theoretically reducing background noise accumulation to O(n). Empirical results on InfiniteBench, RULER, and LongBench show that CRISP outperforms existing sparse methods and can match or exceed exact dense attention, achieving up to a 5.30× speedup at 512k tokens and significant gains on retrieval-heavy tasks.

By Huu Huy Nguyen, Chien Van Nguyen, Franck Dernoncourt, Ryan A. Rossi, Linh Ngo Van, Jieyang Chen, Thien Huu Nguyen
arXiv AI
Aug 19

TileMix: Tile-Centric Mixed-Precision Attention for LLM Inference Acceleration

TileMix is a tile‑centric mixed‑precision attention kernel that routes score‑tile groups within fused dense attention to either FP16 or INT8 computation, using compact bitmasks to decide precision per tile. By partitioning the attention matrix into hardware‑aligned tiles and updating a shared online‑softmax state, TileMix preserves dense token connectivity without requiring training and supports grouped‑query attention, variable‑length batches, and INT8 key/value caches. Benchmarks on LLaMA, Qwen, and Vicuna show that TileMix restores long‑context quality lost with uniform INT8 and improves prefill throughput over FP16, offering a controllable accuracy‑efficiency trade‑off across model families.

By Hanzhi Zhang, Qiao Zhang, Qinglei Cao, Heng Fan, Yan Huang, Kewei Sha, Yunhe Feng