arXiv Machine Learning By Solal Vernier, Ivan Can Arisoy, Merwan Barlier, Bla\v{z} \v{S}krlj

Building a User Foundation Model for the Open Web

Read the original on arXiv Machine Learning →

arXiv:2607. 28019v1 Announce Type: new Abstract: User foundation models have demonstrated strong results in e-commerce and social recommendation, but most industrial deployments assume environments where user identity is stable and persistent.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv Machine Learning
Aug 12

CADET: Context-Conditioned Ads CTR Prediction With a Decoder-Only Transformer

arXiv:2602. 11410v2 Announce Type: replace Abstract: Click-through rate (CTR) prediction is fundamental to online advertising systems.

By David Pardoe, Neil Daftary, Miro Furtado, Aditya Aiyer, Yu Wang, Liuqing Li, Tao Song, Lars Hertel, Young Jin Yun, Senthil Radhakrishnan, Zhiwei Wang, Tommy Li, Khai Tran, Ananth Nagarajan, Ali Naqvi, Yue Zhang, Renpeng Fang, Avi Romascanu, Arjun Kulothungun, Deepak Kumar, Praneeth Boda, Fedor Borisyuk, Ruoyan Wang
arXiv AI
Sep 2

From Language to Behavior: Scaling Sequence Transformers for Industrial Recommendation Ranking with Rec-Native Designs

The paper introduces ReST, a recommendation‑native Transformer scaling framework designed to handle noisy, irregular, and sparsely supervised user behavior sequences in production ranking. ReST employs a dual‑gated attention encoder with rotary positional and temporal embeddings, and a lightweight cross decoder that decouples heavy encoding from fast decoding, enabling efficient compute‑once, decode‑many‑times ranking. Experiments on industrial and public benchmarks show that ReST outperforms traditional Transformer blocks, achieving higher accuracy and consistent scaling across sequence length, depth, and width, and a one‑week online A/B test on a production advertising platform yielded a 1.31% AUC lift and an 11.93% increase in a core revenue metric within a 50 ms P99 latency budget.

By Jie Chen, Xiangqian Yu, Yanchao Lian, Tan Lu, Run Yang, Zhengchun Shang, Xing Wang, Cheng Chen, Ke Hu, Qiang Li, Tianjiu Yin, Xiaobing Liu
Hugging Face Trending Papers
Sep 2

Discriminative World Models for Web Agents

Discriminative World Models for Web Agents proposes a new training objective called predicted‑state matching, which forces a world model to produce representations that can distinguish the true resulting web state from those produced by alternative actions. The authors train these models on a branching dataset from WebArena Go‑Browse, where each decision point includes multiple actions and their outcomes. Experiments show that models trained with predicted‑state matching outperform those trained with standard supervised next‑state prediction on a held‑out benchmark, improve PRM‑style action ranking on WebPRMBench, and enhance end‑to‑end task success on WebArena‑Lite when used for test‑time action selection.

arXiv Machine Learning
Sep 17

Behavioral Fingerprinting and Navigation Prediction in Web Browsing

The study examines two behavioral inference tasks—session-level user identification and next-domain prediction—using large-scale anonymous web browsing traces. Classical and neural models are applied to user identification, while graph-based methods combined with Large Language Models (LLMs) are used for next-domain prediction. Results show that short browsing sessions are highly identifiable and future navigation is highly predictable, with LLM-derived semantic features offering only marginal improvements over structural and sequential models.

By Ralph Elsaghbini, Omran Berjawi, Walid Fahs, Rida Khatoun
arXiv AI
Jun 26

From Clicks to Intent: Cross-Platform Session Embeddings with LLM-Distilled Taxonomy for Financial Services Recommendations

arXiv:2606. 26277v1 Announce Type: cross Abstract: Sequential user behavior modeling is widely adopted in industrial recommender systems; however, significant gaps remain in financial services, where pre-login web interactions and authenticated in-app experiences differ drastically.

By Dianjing Fan, Yao Li, Kyaw Hpone Myint, Dwipam Katariya, Alexandre G. R. Day, Pranab Mohanty, Giri Iyengar