arXiv:2509. 07963v2 Announce Type: replace Abstract: The core component of attention is the scoring function, which transforms the inputs into low-dimensional queries and keys and takes the dot product of each pair.
By Yilun Kuang, Noah Amsel, Sanae Lotfi, Shikai Qiu, Andres Potapczynski, Andrew Gordon Wilson
arXiv:2609.08615v1 Announce Type: new
Abstract: Dimensional attention in learning is often implemented as a globally shared attention vector, where each stimulus dimension corresponds to a single sca...
By Lenard Dome
arXiv:2604. 00757v2 Announce Type: replace-cross Abstract: Large Vision Language Models show impressive performance across image and video understanding tasks, yet their computational cost grows rapidly with the number of visual tokens.
By Dong-Jae Lee, Sunghyun Baek, Junmo Kim
arXiv:2504. 18455v2 Announce Type: replace-cross Abstract: We study distributed multiview representation learning, a problem in which $K$ clients each observe a distinct but possibly statistically correlated view.
By Milad Sefidgaran, Piotr Krasnowski, Abdellatif Zaidi
arXiv:2606. 01948v1 Announce Type: cross Abstract: The growing popularity of group activities has increased the need for methods that provide recommendations to groups of users given their individual preferences.
By Mubaraka Sani Ibrahim, Lehel Csat\'o, Isah Charles Saidu
arXiv:2607. 24025v1 Announce Type: cross Abstract: Transformer architectures have achieved remarkable success across diverse domains; however, directly applying their standard self-attention mechanism to recommendation often yields suboptimal performance, sometimes even trailing behind well-designed simple recommendation models.
By Yu Cui, Yi Xu, Jiahao Wang, Hao Zhang, Yu Zhang, Xiaoyi Zeng, Can Wang, Jinxin Hu, Jiawei Chen
In federated averaging, local objectives can admit multiple optimal heads, making the aggregate depend on which heads clients return. We study this ambiguity in federated multivariate regression with...
arXiv:2607. 20214v1 Announce Type: cross Abstract: The quadratic $N\times N$ attention score matrix remains a central obstacle to extending Transformers to longer input lengths.
By Mahdi Heidari, Mohammad Mahdi Rahimi, Jaekyun Moon
arXiv:2504. 17768v3 Announce Type: replace-cross Abstract: Sparse attention offers a promising strategy to extend long-context capabilities in Transformer LLMs, yet its efficiency-accuracy trade-offs remain unclear due to the lack of comprehensive evaluation.
By Piotr Nawrot, Robert Li, Renjie Huang, Sebastian Ruder, Kelly Marchisio, Edoardo M. Ponti
arXiv:2607. 02344v1 Announce Type: cross Abstract: Transformer architectures have shown strong potential in time series forecasting, where multi-head self-attention is widely used to capture temporal dependencies across historical timestamps.
By Dezheng Wang, Tong Chen, Wei Yuan, Congyan Chen, Shihua Li, Hongzhi Yin
arXiv:2607. 02681v1 Announce Type: cross Abstract: Integrating information across related tasks can improve estimation and prediction in transfer, multi-task, and federated learning, but contamination and heterogeneity make robust borrowing challenging.
By Ye Tian, Mengchu Li, Marco Avella Medina
arXiv:2608.21443v1 Announce Type: new
Abstract: Estimating interpretable conditional-dependence structures from multimodal visual-linguistic features remains largely unexplored. We propose CM-GLasso...
By Fei Wang, Yutong Zhang, Yang Ye, Jinxian Chen, Wang Wenshuai, Xiong Wang