arXiv AI

Action-Aware Generative Sequence Modeling for Short Video Recommendation

arXiv:2604. 25834v2 Announce Type: replace Abstract: With the rapid development of the Internet, users have increasingly higher expectations for the recommendation accuracy of online content consumption platforms.

arXiv AI
Sep 10

SequenceO1: End-to-End Ultra-Long (100K) Sequence Modeling in Recommendation with Low-Rank Caching

SequenceO1 is an end‑to‑end framework that enables ultra‑long (up to 100K interactions) sequence modeling for recommendation systems. It compresses raw user histories into a fixed‑size sketch using Sketch Attention and then models short‑term and long‑term interests with Target‑to‑History Cross Attention. The system incorporates low‑rank caching, batching, pipeline lift, and a FlashSA kernel to keep training and inference efficient, achieving consistent offline and online performance gains when deployed at full traffic on Douyin.

By Lin Guan, Jia-Qi Yang, Zhishan Zhao, Jiaqi Huang, Hangyu Wang, Longbin Li, Beichuan Zhang, Haonan Jiang, Jinan Ni, Xiangyu Fan, Xiaowen Li, Ziyao Ren, Yuhang Qi, Xiaolong Zhu, Xuanyuan Luo, Qiwei Chen, Yi Cheng, Lele Yu
arXiv AI
Jun 9

Beyond Item IDs: Scaling Short-Form-Video Recommendation via Semantic-Native Long Sequence Modeling

arXiv:2606. 07546v1 Announce Type: cross Abstract: Capturing user interests across extensive watch histories is critical for short-form video recommendation, yet scaling sequence length is limited by two bottlenecks: the semantic sparsity of atomic Video IDs and the quadratic computational complexity of Transformers.

By Ruixiao Sun, Diego Uribe Mora, Zhimeng Jiang, Yuanzhen Lin, Jiarui Wang, Yuening Li, Danfeng Guo, Zhizhong Chen, Chuan He, Liang Liu
arXiv AI
Sep 21

Dual-Interest Sequential Product Recommendation With Multi-Granular SSM

The paper introduces DSRec, a dual‑interest sequential recommendation model that separates item representations into long‑term and short‑term semantic contexts. Long‑term embeddings capture stable preferences through historical aggregation, while short‑term embeddings focus on local session intent modulated by inter‑click time intervals. Each branch is processed by a distinct State Space Model— a full‑sequence Mamba for long‑term modeling and a time‑modulated SSM for short‑term dynamics— and a residual cross‑fusion mechanism aligns the two granularities while preserving their independence. Experiments on public benchmarks show that DSRec outperforms state‑of‑the‑art methods.

By Shuiying Liao, P. Y. Mok
arXiv AI
Jun 30

CMSL: Constructive Multi-Sequence Learning for Recommendation Systems

arXiv:2606. 28533v1 Announce Type: cross Abstract: Sequence learning has emerged as the promising paradigm in recommendation systems, surpassing traditional Deep Learning Recommendation Models (DLRM) by capturing the temporal nuances of user behavior.

By Zikun Cui, Renzhi Wu, Junjie Yang, Li Sheng, Jijie Wei, Linfeng Liu, Tai Guo, Tao Jia, Xiaodong Wang, Hong Li, Li Yu, Sri Reddy, Hong Yan
arXiv Machine Learning
Jun 25

TokenMinds: Pretrained User Tokens and Embeddings for User Understanding in Large Recommender Systems

arXiv:2606. 25147v1 Announce Type: cross Abstract: User modeling in industrial recommender systems typically produces dense embeddings, which suffer from representational constraints inherent to fixed-dimensional vectors.

By Qingyun Liu, Bo Yan, Yang Liu, Yuji Roh, Ekansh Sharma, Likang Yin, Emma Olowo, Min-hsuan Tsai, Yuxuan Li, Diego Uribe, Saksham Aggarwal, Siqi Wu, Yuan Hao, Vikas Kedigehalli, Lukasz Heldt, Lichan Hong, Li Wei, Xinyang Yi
arXiv AI
2d ago

NextMe-800: Anticipating Personal Behavior from Months of Egocentric Video

NextMe-800 is an approximately 800‑hour first‑person video dataset collected from a single volunteer over 126 days, featuring 1 Hz images, gaze, and audio. The data are captioned at five hierarchical abstraction levels—from atomic actions to major activities—enabling personalized action anticipation as an open‑vocabulary K‑step sequence prediction task. The authors also introduce NextAct, a 1,500‑point benchmark that combines NextMe‑800 with the multi‑person EgoLife dataset, and evaluate models using an embedding‑based soft edit distance to assess how well personal behavior can be anticipated across abstraction levels and prediction horizons.

By Zhaoxu Meng, Yiming Sun, Mingyuan Gao, Jiachang Zhang, Zhuhan Dai, Yipeng Du, Zheng Lian, Jian-Qiao Zhu
arXiv AI
Aug 19

M3TR: Temporal Retrieval Enhanced Multi-Modal Micro-video Popularity Prediction

M3TR is a temporal retrieval‑enhanced multi‑modal framework for predicting micro‑video popularity. It introduces a Mamba‑Hawkes Process module to model user feedback as self‑exciting events, capturing long‑range temporal dependencies. A temporal‑aware retrieval engine then identifies historically relevant videos by combining multi‑modal content similarity with popularity trajectory similarity, augmenting the target video’s features for improved prediction accuracy.

By Jiacheng Lu, Weijian Wang, Mingyuan Xiao, Yang Hua, Tao Song, Bo Peng, Cheng Hua, Haibing Guan
arXiv AI
Jul 21

WHALE: A Scalable Unified Model for Recommendation with Wukong-HSTU Architecture

arXiv:2607. 17017v1 Announce Type: cross Abstract: As scalability becomes increasingly important in recommendation modeling, recent architectures have advanced the modeling of two broad sources of ranking signals along separate paths: non-sequence features, including user, item, context, and cross features; and sequence features from user behavior histories.

By Renqin Cai, Dawei Sun, Yuanjun Yao, Zhiyong Wang, Velvin Fu, Maggie Zhuang, Yu Shi, Zhongnan Fang, Xuan Cao, Jing Qian, Rui Li