arXiv AI

From Clicks to Intent: Cross-Platform Session Embeddings with LLM-Distilled Taxonomy for Financial Services Recommendations

arXiv:2606. 26277v1 Announce Type: cross Abstract: Sequential user behavior modeling is widely adopted in industrial recommender systems; however, significant gaps remain in financial services, where pre-login web interactions and authenticated in-app experiences differ drastically.

arXiv AI
Sep 2

User Representation via Cross Multi-source Behavior Pre-training for Mobile Games

The paper introduces CM-PTM, a Cross Multi-source Behavior Pre-Training Model designed to learn mobile game user representations from device-level behavioral logs. It uses hierarchical cascaded mask‑then‑predict tasks to first identify the source of the next behavior and then refine predictions at the app‑action level, thereby modeling cross‑source dependencies and fine‑grained dynamics. Experiments on large real‑world mobile datasets show that CM‑PTM captures users’ endogenous interests and improves performance on downstream mobile game recommendation tasks.

By Chengqi Yang, Yiran Qiao, Feng Liu, Xingyu Lou, Zijun Zhou, Xiaoyun Mo, Changwang Zhang, Jiayuan Xu, Jun Wang, Xiang Ao
arXiv Machine Learning
Sep 3

MISApp: Multi-Hop Intent-Aware Session Graph Learning for Next App Prediction

MISApp is a profile‑free framework that predicts the next mobile app a user will launch by learning multi‑hop session graphs. It captures transition dependencies across different structural ranges, incorporates temporal context and spatial categorization, and models intent evolution from recent interactions. Experiments on two real‑world datasets show MISApp outperforms baselines in both standard and cold‑start settings while remaining efficient, and analyses reveal that multi‑hop relations provide higher‑order predictive signals and interpretable attention weights.

By Yunchi Yang, Longlong Li, Jianliang Wu, Cunquan Qu
arXiv Machine Learning
Sep 17

Behavioral Fingerprinting and Navigation Prediction in Web Browsing

The study examines two behavioral inference tasks—session-level user identification and next-domain prediction—using large-scale anonymous web browsing traces. Classical and neural models are applied to user identification, while graph-based methods combined with Large Language Models (LLMs) are used for next-domain prediction. Results show that short browsing sessions are highly identifiable and future navigation is highly predictable, with LLM-derived semantic features offering only marginal improvements over structural and sequential models.

By Ralph Elsaghbini, Omran Berjawi, Walid Fahs, Rida Khatoun
arXiv Machine Learning
Jun 25

TokenMinds: Pretrained User Tokens and Embeddings for User Understanding in Large Recommender Systems

arXiv:2606. 25147v1 Announce Type: cross Abstract: User modeling in industrial recommender systems typically produces dense embeddings, which suffer from representational constraints inherent to fixed-dimensional vectors.

By Qingyun Liu, Bo Yan, Yang Liu, Yuji Roh, Ekansh Sharma, Likang Yin, Emma Olowo, Min-hsuan Tsai, Yuxuan Li, Diego Uribe, Saksham Aggarwal, Siqi Wu, Yuan Hao, Vikas Kedigehalli, Lukasz Heldt, Lichan Hong, Li Wei, Xinyang Yi
arXiv Machine Learning
Aug 18

SAGA: Structure-Attended Generative Action Embedding Model that encodes Multi-Surface User Action Sequences

arXiv:2608. 15429v1 Announce Type: new Abstract: Prior embedding models for sequential recommendation typically operate within a homogeneous action space, limiting their ability to capture cross-surface behavioral signals spanning distinct behavioral domains.

By Tsz Fung Pang, Po Jen Chen, Nimish Ronghe, Farhad Farahani, Bo Zhang
arXiv Computation and Language
Sep 3

Beyond Retrieval: Learning Compact User Representations for Scalable LLM Personalization

The paper introduces TAP-PER, a prefix‑based framework that learns compact user representations for large language model personalization. By encoding user preferences into lightweight prefix embeddings and incorporating temporal signals, TAP‑PER avoids the need for heavy per‑user adapters or prompt‑serialized histories. Experiments on six LaMP tasks show that TAP‑PER outperforms both prompt‑based and model‑based baselines while using far fewer per‑user parameters, enabling scalable personalization at large user scales.

By Heng Cao, Fan Zhang, Jian Yao, Yujie Zheng, Changlin Zhao, Lu Hao, Yuxuan Wei, Wangze Ni, Huaiyu Fu, Yuqian Sun, Xuyan Mo