arXiv AI

QueryFormer: Winning Solution for KDD Cup 2026 Tencent UniRec Challenge

QueryFormer is a unified architecture designed for post‑click conversion rate prediction, addressing both feature interactions and sequential user behaviors. It introduces a stackable field–sequence block that generates query tokens via cross‑attention and packs sequence queries into shared‑parameter attention, improving efficiency and accuracy. The model won first place in the KDD Cup 2026 Tencent UniRec Challenge Industrial Track with an AUC of 0.83254, and scaling studies show that increasing view width slightly boosts validation AUC while maintaining low latency.

arXiv Machine Learning
Sep 21

Elastic Threshold Attention: Learned Contextual Sparsity for Long-Context Decoding

Elastic Threshold Attention (ETA) is a trainable attention mechanism that dynamically predicts contextual thresholds from query representations, enabling selective pruning of KV cache tokens during long‑context decoding. By multiplicatively suppressing sub‑threshold logits during training, ETA avoids representation collapse and eliminates localized attention sinks, allowing a 1.45B model to match dense attention performance at roughly 85% training sparsity and 38% active decode density. At inference, a custom Triton kernel achieves up to 2.5× faster decoding on sequences up to 512K tokens, and an offline calibration step can further reduce compute by 27% by freezing per‑head thresholds.

By Themistoklis Haris, Henry Li, Maryam Karimzadehgan
arXiv Machine Learning
Jul 31

ROCS: Request-Oriented Compute Sharing for Efficient Large-Scale Recommendation

arXiv:2607. 27744v1 Announce Type: new Abstract: Modern recommendation models gain prediction quality by scaling feature-interaction and sequence modules, but production cost constraints cap how far systems can scale.

By Yuxin Chen, Liang Luo, Buyun Zhang, Jian Jiao, Boda Li, Haoyu Wang, Tongyi Tang, Ao Cai, Zijian Shen, Zhengkai Zhang, Wenyi Xie, Ryan Dick, Han Liu, Neng Shi, Bin Yu, Jianbo Xiao, Shuyao Bi, Hongtao Yu, Yuanwei Fang, Zhuoran Zhao, Sijia Chen, Yang Chen, Shuqi Yang, Qianru Li, Zikun Liu, Wei Ling, Sihan Zeng, Longhao Jin, Jiaxin Lu, Yinbin Ma, Jiawei Li, Yichen Ruan, Yong Ler Lee, Birmingham Guan, Zijian Li, Jianbo Sun, Zhengyu Zhang, Zeliang Chen, Xiaohan Wei, Yuchen Hao, GP Musumeci, Venkatesh Ranganathan, Yantao Yao, Chunqiang Tang, Wenlin Chen, Santanu Kolay, Ellie Dingqiao Wen
arXiv Machine Learning
Sep 18

Post-Boundary Bridge: Must Local Attention Go Global Between Global Layers?

The paper introduces Post-Boundary Bridge (PBB), a hybrid transformer architecture that keeps causal attention within blocks while adding direct connections across block boundaries. PBB focuses on within-block modeling and nearby exchange, delegating long-range communication to full-attention layers. Experiments on dense and mixture-of-experts models ranging from 205 million to 2.07 billion parameters show that PBB hybrids maintain near-full perplexity, competitive downstream performance, and improved source retrieval, while Flash-PBB achieves 1.82× faster decoding with half the local key‑value cache compared to Flash‑SWA.

By Zhibo Yang
arXiv AI
Sep 25

Near-Oracle KV Selection via Pre-hoc Sparsity for Long-Context Inference

The paper introduces Pre-hoc Sparsity (PrHS), a method that selects key-value (KV) cache entries before attention scoring to avoid posterior bias in large language model inference. By bounding mutual‑information loss through the dropped attention mass, PrHS offers explicit accuracy control and implements three orthogonal selectors across time, depth, and layer. Experiments on LLaMA and Mistral models show that PrHS cuts retrieval overhead by over 90%, achieves higher sparsity than HShare, and delivers significant speedups and reduced FLOPs on NVIDIA A100 GPUs while maintaining near‑dense accuracy.

By Yifei Gao, Lei Wang, Rong-Cheng Tu, Qixin Zhang, Jun Cheng, Dacheng Tao