arXiv AI By Keito Kozaki, Keigo Sakurai, Ren Togo, Takahiro Ogawa, Miki Haseyama

Residual Dominance as a Structural Account of Last-Item Reliance in Causal Self-Attention Recommenders

Read the original on arXiv AI →

arXiv:2608. 14021v1 Announce Type: new Abstract: Transformer-based sequential recommenders with causal self-attention often rely heavily on the most recent interaction at inference time, but how this behavior is structurally expressed in the representation used for prediction remains unclear.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Machine Learning
Jul 28

SpecFormer: Mitigating Embedding and Attention Collapse via Spectral-Aware Transformer for Recommendation

arXiv:2607. 24025v1 Announce Type: cross Abstract: Transformer architectures have achieved remarkable success across diverse domains; however, directly applying their standard self-attention mechanism to recommendation often yields suboptimal performance, sometimes even trailing behind well-designed simple recommendation models.

By Yu Cui, Yi Xu, Jiahao Wang, Hao Zhang, Yu Zhang, Xiaoyi Zeng, Can Wang, Jinxin Hu, Jiawei Chen
arXiv AI
Sep 12

Relevance Is Not Permission: Localizing and Controlling Metric-Facing Attention Contributions

The paper introduces Warrant, a method that locates and controls the contributions of attention mechanisms to model metrics. Warrant exposes the item‑wise contribution path to the reported metric and applies query‑conditioned permission on that path. Experiments on multiple datasets show that Warrant improves primary metrics in most comparisons, reveals a weak correlation between attention and prediction utility, and demonstrates that learned permission can recover evidence ranking while suppressing distractors.

By Minwoo Yu, Young-guk Ha
arXiv AI
Sep 24

Warranted Attention: Learning What to Pass from Attention to Prediction

The paper introduces Warrant, a method that learns to gate attention-derived item contributions before they are aggregated for prediction. Unlike traditional attention, which assumes relevance guarantees usefulness, Warrant applies learned, item‑wise permissions to control both the relative allocation and total transmission mass. Experiments on CyGNet and HotpotQA datasets show that ungated attention paths degrade performance, while Warrant’s selective gating recovers or improves metrics such as MRR and reduces unsupported selections.

By Minwoo Yu, Young-guk Ha
arXiv Machine Learning
Sep 25

Beyond Pairwise Attention: Higher-Order Modular Attention for Efficient Sequence Learning

The paper introduces Higher-Order Modular Attention (HOMA), a new attention mechanism that combines standard pairwise self‑attention with an explicit triadic attention pathway. HOMA uses overlapping blocks, local windows, and a low‑rank projection to make triadic interactions tractable. Experiments on controlled PARITY and MATCH3 tasks, as well as TAPE benchmarks, show that HOMA matches or outperforms matched pairwise and purely triadic baselines, especially when dependencies extend beyond triadic order, and it often converges faster and uses parameters more efficiently.

By Shirin Amiraslani, Xin Gao
arXiv Machine Learning
Sep 24

A Systematic Benchmark of Explainable Methods for Temporal Attribution in Sequential Recommendation Systems

The paper introduces a systematic benchmark for evaluating explainable methods that attribute temporal interactions in sequential recommendation systems. Using a dual-model masking metric, it assesses ten XAI techniques across CNN, Transformer, SASRec, and BERT4Rec backbones on KuaiRand and MovieLens datasets, revealing that gradient-based methods like GradientSHAP and Integrated Gradients are the most faithful and robust. It also finds that raw attention weights are unreliable, while gradient-weighted attention works better on short sequences but degrades on longer horizons, and that faithful methods capture genuine task structure rather than recency or popularity bias.

By Akash Pandey, Kanisha Shah, Addrish Roy, Dwipam Katariya, Hongyangyang Shi, Amanda Ding, Kalanand Mishra, Pranab Mohanty
arXiv AI
Sep 3

CaST-POI: Candidate-Conditioned Spatiotemporal Modeling for Next POI Recommendation

CaST-POI is a next‑point‑of‑interest recommender that conditions the user representation on each candidate location by adding bucketised biases for visit recency and geographic distance to the candidate. Unlike prior models that treat all candidates uniformly, CaST-POI lets each candidate read the trajectory with different attention weights grounded in real distance. Experiments on NYC, TKY, and CA datasets show significant MRR gains over seven baselines, with ablation revealing the importance of the revisit gate and spatial bias.

By Zhenyu Yu, Chunlei Meng, Yangchen Zeng, Mohd Yamani Idna Idris, Jihong Guan, Shuigeng Zhou