Multi-modal recommenders fuse collaborative signals with item modalities such as text, images, and audio, but the usefulness of each drifts over time and at different rates. For example, chocolate purchases typically guided by textual ingredient cues can shift toward visual packaging and ambient audio around Valentine's Day.
arXiv:2606. 01670v1 Announce Type: cross Abstract: Recently, Generative Recommenders (GRs) have emerged as a transformative recommendation paradigm by replacing traditional item IDs with semantic indices (SIDs).
By Bangguo Zhu, Peng Huo, Yuanbo Zhao, Zhicheng Du, Jun Yin, Senzhang Wang
Recently, Generative Recommenders (GRs) have emerged as a transformative recommendation paradigm by replacing traditional item IDs with semantic indices (SIDs). Owing to the exceptional generative capabilities of diffusion models, a few pioneering works explore developing GRs with diffusion architectures as the backbone.
arXiv:2608.23400v1 Announce Type: cross
Abstract: Discrete Diffusion Models (DDMs) have recently been introduced to recommendation systems, modeling user history as a token generation process via ite...
By Jiaqi Wang, Tianying Liu, Heng Chang, Jihong Guan, Wengen Li, Shuigeng Zhou
Discrete Diffusion Models (DDMs) have recently been introduced to recommendation systems, modeling user history as a token generation process via iterative denoising. However, while effective at captu...
arXiv:2608. 10240v1 Announce Type: cross Abstract: Multi-modal sequential recommenders assume every item carries every modality, but real product catalogs often miss images or text, and a model trained on complete data loses much of its recommendation accuracy when a modality is unavailable at serving time.
By Guanqun Yang, Wenlong Zhang
arXiv:2604. 02183v3 Announce Type: replace Abstract: Multimodal recommendation systems (MRS) jointly model user-item interaction graphs and rich item content, but this tight coupling makes user data difficult to remove once learned.
By Zhanting Zhou, KaHou Tam, Ziqiang Zheng, Zeyu Ma, Yang Yang
arXiv:2602.20723v3 Announce Type: replace
Abstract: Multimodal recommenders combine collaborative behavior with visual and textual item evidence, whose usefulness varies across user-item interactions...
By Ji Dai, Quan Fang, DeSheng Cai
Positional encoding is a fundamental component of Transformer-based generative recommendation models, where user histories are modeled as autoregressive item sequences. Most positional encoding methods are inherited from natural language processing and mainly represent discrete item order.
arXiv:2608. 16274v1 Announce Type: cross Abstract: Positional encoding is a fundamental component of Transformer-based generative recommendation models, where user histories are modeled as autoregressive item sequences.
By Pengfei Jia, Jingjian Wang, Jingmao Li, Ge Zhang, Feng Shi
arXiv:2606. 06225v1 Announce Type: cross Abstract: Collaborative filtering and graph-based recommendation models are highly effective because they leverage observed user interactions, but this dependence creates a fundamental cold-start challenge when newly added content has no interaction history.
By Anh Truong, John Trenkle, Yuanbo Chen, Honghong Zhao, Abdullah Alchihabi, Effy Fang, Michael Tamir
arXiv:2606. 01352v1 Announce Type: new Abstract: Watch time has emerged as a pivotal metric for optimizing deep user engagement in short-video recommender systems.
By Hongxu Ma, Han Zhou, Chenghou Jin, Jie Zhang, Xiaoyu Yang, Chunjie Chen, Jihong Guan, Shuigeng Zhou