arXiv AI By Yuanzhe Zhou, Zhaoyang Zeng

QueryFormer: Winning Solution for KDD Cup 2026 Tencent UniRec Challenge

Read the original on arXiv AI →

QueryFormer is a unified architecture designed for post‑click conversion rate prediction, addressing both feature interactions and sequential user behaviors. It introduces a stackable field–sequence block that generates query tokens via cross‑attention and packs sequence queries into shared‑parameter attention, improving efficiency and accuracy. The model won first place in the KDD Cup 2026 Tencent UniRec Challenge Industrial Track with an AUC of 0.83254, and scaling studies show that increasing view width slightly boosts validation AUC while maintaining low latency.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Machine Learning
Sep 21

Elastic Threshold Attention: Learned Contextual Sparsity for Long-Context Decoding

Elastic Threshold Attention (ETA) is a trainable attention mechanism that dynamically predicts contextual thresholds from query representations, enabling selective pruning of KV cache tokens during long‑context decoding. By multiplicatively suppressing sub‑threshold logits during training, ETA avoids representation collapse and eliminates localized attention sinks, allowing a 1.45B model to match dense attention performance at roughly 85% training sparsity and 38% active decode density. At inference, a custom Triton kernel achieves up to 2.5× faster decoding on sequences up to 512K tokens, and an offline calibration step can further reduce compute by 27% by freezing per‑head thresholds.

By Themistoklis Haris, Henry Li, Maryam Karimzadehgan