arXiv AI

From Feature Interaction to Feature Transport - A Unified Block for Scalable Recommendation Models

The paper introduces CRAFT, a Contextual Residual Adaptive Feature Transport block that treats unified recommendation models as a discrete context‑conditioned representation evolution process. By summarizing non‑sequential features into a reliability‑aware contextual field, CRAFT generates residual displacement and memory‑preserving signals to control intent and sequence representations. Experiments on the TAAC2026 competition show CRAFT achieving a test AUC of 0.838090, surpassing the previous best, and further improvements with deeper or wider models.

arXiv AI
Jul 21

WHALE: A Scalable Unified Model for Recommendation with Wukong-HSTU Architecture

arXiv:2607. 17017v1 Announce Type: cross Abstract: As scalability becomes increasingly important in recommendation modeling, recent architectures have advanced the modeling of two broad sources of ranking signals along separate paths: non-sequence features, including user, item, context, and cross features; and sequence features from user behavior histories.

By Renqin Cai, Dawei Sun, Yuanjun Yao, Zhiyong Wang, Velvin Fu, Maggie Zhuang, Yu Shi, Zhongnan Fang, Xuan Cao, Jing Qian, Rui Li
arXiv Machine Learning
Aug 18

SAGA: Structure-Attended Generative Action Embedding Model that encodes Multi-Surface User Action Sequences

arXiv:2608. 15429v1 Announce Type: new Abstract: Prior embedding models for sequential recommendation typically operate within a homogeneous action space, limiting their ability to capture cross-surface behavioral signals spanning distinct behavioral domains.

By Tsz Fung Pang, Po Jen Chen, Nimish Ronghe, Farhad Farahani, Bo Zhang
arXiv AI
Sep 21

Dual-Interest Sequential Product Recommendation With Multi-Granular SSM

The paper introduces DSRec, a dual‑interest sequential recommendation model that separates item representations into long‑term and short‑term semantic contexts. Long‑term embeddings capture stable preferences through historical aggregation, while short‑term embeddings focus on local session intent modulated by inter‑click time intervals. Each branch is processed by a distinct State Space Model— a full‑sequence Mamba for long‑term modeling and a time‑modulated SSM for short‑term dynamics— and a residual cross‑fusion mechanism aligns the two granularities while preserving their independence. Experiments on public benchmarks show that DSRec outperforms state‑of‑the‑art methods.

By Shuiying Liao, P. Y. Mok
arXiv Machine Learning
Jun 25

TokenMinds: Pretrained User Tokens and Embeddings for User Understanding in Large Recommender Systems

arXiv:2606. 25147v1 Announce Type: cross Abstract: User modeling in industrial recommender systems typically produces dense embeddings, which suffer from representational constraints inherent to fixed-dimensional vectors.

By Qingyun Liu, Bo Yan, Yang Liu, Yuji Roh, Ekansh Sharma, Likang Yin, Emma Olowo, Min-hsuan Tsai, Yuxuan Li, Diego Uribe, Saksham Aggarwal, Siqi Wu, Yuan Hao, Vikas Kedigehalli, Lukasz Heldt, Lichan Hong, Li Wei, Xinyang Yi
arXiv Machine Learning
Jul 14

Tokenizing Numerical and Embedding Features for LLM RecSys

arXiv:2607. 10016v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly used as backbone architectures for recommender systems because of their strong sequence modeling and representation learning capabilities.

By Zhe Xu, Ankit Peshin, Chiyu Zhang, Feng Qi, Johnson Lui, Anil Ramakrishna, Justin Johnson, Carl Hu, Kaushik Rangadurai, Luke Simon
arXiv AI
Sep 10

SequenceO1: End-to-End Ultra-Long (100K) Sequence Modeling in Recommendation with Low-Rank Caching

SequenceO1 is an end‑to‑end framework that enables ultra‑long (up to 100K interactions) sequence modeling for recommendation systems. It compresses raw user histories into a fixed‑size sketch using Sketch Attention and then models short‑term and long‑term interests with Target‑to‑History Cross Attention. The system incorporates low‑rank caching, batching, pipeline lift, and a FlashSA kernel to keep training and inference efficient, achieving consistent offline and online performance gains when deployed at full traffic on Douyin.

By Lin Guan, Jia-Qi Yang, Zhishan Zhao, Jiaqi Huang, Hangyu Wang, Longbin Li, Beichuan Zhang, Haonan Jiang, Jinan Ni, Xiangyu Fan, Xiaowen Li, Ziyao Ren, Yuhang Qi, Xiaolong Zhu, Xuanyuan Luo, Qiwei Chen, Yi Cheng, Lele Yu
arXiv Machine Learning
1d ago

Action-On-Item Preference Flow: A Shared Event Schema for Predictive and Generative Personalization

The paper introduces an action‑on‑item schema that pairs interaction roles with content embeddings, enabling a shared update mechanism across different user history types such as movies, news, and dialogue. It demonstrates theoretical properties like invariance to relabeling and bounded state changes, and presents the Multi‑Timescale State Hypothesis (MTSH) implemented in PerTIDE. Experiments on PENS, MovieLens, and MIND datasets show that a frozen source‑trained core outperforms random baselines and that PerTIDE achieves significant MRR gains over comparable models.

By Parthiv Chatterjee, Kashish Kanjaria, Vashisth Purani, Sourish Dasgupta, Tanmoy Chakraborty