arXiv:2607. 24025v1 Announce Type: cross Abstract: Transformer architectures have achieved remarkable success across diverse domains; however, directly applying their standard self-attention mechanism to recommendation often yields suboptimal performance, sometimes even trailing behind well-designed simple recommendation models.
By Yu Cui, Yi Xu, Jiahao Wang, Hao Zhang, Yu Zhang, Xiaoyi Zeng, Can Wang, Jinxin Hu, Jiawei Chen
The paper introduces Warrant, a method that locates and controls the contributions of attention mechanisms to model metrics. Warrant exposes the item‑wise contribution path to the reported metric and applies query‑conditioned permission on that path. Experiments on multiple datasets show that Warrant improves primary metrics in most comparisons, reveals a weak correlation between attention and prediction utility, and demonstrates that learned permission can recover evidence ranking while suppressing distractors.
By Minwoo Yu, Young-guk Ha
The paper introduces Warrant, a method that learns to gate attention-derived item contributions before they are aggregated for prediction. Unlike traditional attention, which assumes relevance guarantees usefulness, Warrant applies learned, item‑wise permissions to control both the relative allocation and total transmission mass. Experiments on CyGNet and HotpotQA datasets show that ungated attention paths degrade performance, while Warrant’s selective gating recovers or improves metrics such as MRR and reduces unsupported selections.
By Minwoo Yu, Young-guk Ha
The paper introduces Higher-Order Modular Attention (HOMA), a new attention mechanism that combines standard pairwise self‑attention with an explicit triadic attention pathway. HOMA uses overlapping blocks, local windows, and a low‑rank projection to make triadic interactions tractable. Experiments on controlled PARITY and MATCH3 tasks, as well as TAPE benchmarks, show that HOMA matches or outperforms matched pairwise and purely triadic baselines, especially when dependencies extend beyond triadic order, and it often converges faster and uses parameters more efficiently.
By Shirin Amiraslani, Xin Gao
The paper introduces a systematic benchmark for evaluating explainable methods that attribute temporal interactions in sequential recommendation systems. Using a dual-model masking metric, it assesses ten XAI techniques across CNN, Transformer, SASRec, and BERT4Rec backbones on KuaiRand and MovieLens datasets, revealing that gradient-based methods like GradientSHAP and Integrated Gradients are the most faithful and robust. It also finds that raw attention weights are unreliable, while gradient-weighted attention works better on short sequences but degrades on longer horizons, and that faithful methods capture genuine task structure rather than recency or popularity bias.
By Akash Pandey, Kanisha Shah, Addrish Roy, Dwipam Katariya, Hongyangyang Shi, Amanda Ding, Kalanand Mishra, Pranab Mohanty
CaST-POI is a next‑point‑of‑interest recommender that conditions the user representation on each candidate location by adding bucketised biases for visit recency and geographic distance to the candidate. Unlike prior models that treat all candidates uniformly, CaST-POI lets each candidate read the trajectory with different attention weights grounded in real distance. Experiments on NYC, TKY, and CA datasets show significant MRR gains over seven baselines, with ablation revealing the importance of the revisit gate and spatial bias.
By Zhenyu Yu, Chunlei Meng, Yangchen Zeng, Mohd Yamani Idna Idris, Jihong Guan, Shuigeng Zhou
arXiv:2511. 21095v2 Announce Type: replace Abstract: Early Stage Ranking (ESR) in large-scale recommendation systems is dominated by ''user--item decoupling'' Two Tower architectures, which scale efficiently but cannot capture fine-grained, target-aware user--item interactions directly.
By Juhee Hong, Meng Liu, Shengzhi Wang, Jin Zhou, Xiaoheng Mao, Zhao Zhu, Ruochen Liu, Huihui Cheng, Leon Gao, Christopher Leung, Chandra Mouli Sekar, Yijia Liu, Boyang Yu, Tuan Trieu, Dawei Sun, Jeet Kanjani, Rui Li, Jing Qian, Xuan Cao, Minjie Fan, Mingze Gao
arXiv:2609.10092v1 Announce Type: cross
Abstract: Large language models (LLMs) increasingly act as research agents, yet their ability to track shifts in research attention is difficult to evaluate be...
By Yingqian Wu, Jingcong Liang, Siyuan Wang, Zhenfei Yin, Philip Torr, Junchi Yu, Zhongyu Wei
arXiv:2606. 07604v1 Announce Type: cross Abstract: Analyzing attention weights has become a standard approach for interpreting the information flow of Large Language Models (LLMs).
By Harry Jake Cunningham, Nicola Muca Cirone
arXiv:2505. 15548v2 Announce Type: replace Abstract: Autoregressive transformer language models frequently exhibit training instability when trained on long sequences, particularly under low-precision arithmetic.
By Suvadeep Hajra
InfoMamba is an attention‑free hybrid model that combines a minimal‑bandwidth global interface with a selective recurrent stream. The architecture replaces token‑level self‑attention with a concept bottleneck linear filtering layer and integrates it via an information‑maximizing fusion (IMF) that injects global context into the state‑space dynamics. Experiments across classification, dense prediction, and non‑vision tasks show that InfoMamba outperforms strong Transformer and SSM baselines while maintaining near‑linear scaling and competitive accuracy‑efficiency trade‑offs.
By Youjin Wang, Jiaqiao Zhao, Rong Fu, Run Zhou, Ruizhe Zhang, Jiani Liang, Suisuai Cao, Feng Zhou
arXiv:2509. 07963v2 Announce Type: replace Abstract: The core component of attention is the scoring function, which transforms the inputs into low-dimensional queries and keys and takes the dot product of each pair.
By Yilun Kuang, Noah Amsel, Sanae Lotfi, Shikai Qiu, Andres Potapczynski, Andrew Gordon Wilson