arXiv Machine Learning

Scaling Laws for Behavioral Foundation Models over User Event Sequences

arXiv:2606. 05257v1 Announce Type: new Abstract: Foundation models are increasingly trained on sequences of user actions in recommendation, payments, fraud, and commerce, but these models still lack the kind of compute calibration that scaling laws provide for language models.

arXiv AI
Jun 9

Item Response Scaling Laws: A Measurement Theory Approach for Efficient and Generalizable Neural Scaling Estimation

arXiv:2606. 07616v1 Announce Type: cross Abstract: Scaling laws provide a fundamental framework for understanding the performance of Language Models (LMs), yet deriving them requires prohibitively expensive evaluations across thousands of checkpoints or millions of inference samples.

By Sang Truong, Yuheng Tu, Rylan Schaeffer, Sanmi Koyejo
arXiv AI
Sep 4

Learning What Not to Forget: Long-Horizon Agent Memory from a Few Kilobytes of Learning

The paper introduces LRE (Learned Relevance Eviction), a lightweight, CPU‑only, language‑model‑free scorer that learns which parts of an agent’s interaction history are task‑critical and preserves them verbatim. In experiments, LRE matches or surpasses baseline eviction policies on accuracy‑cost trade‑offs, recovers 93% of full‑history accuracy, reduces worst‑case prompt size by 52%, and outperforms dense and token‑pruning encoders in conversational memory while being 295–1569× smaller. The method also achieves superior budgeted answer quality on LoCoMo reading and can be trained annotation‑free, recovering 95% of supervised scorer performance.

By Nusrat Jahan Lia, Aritra Mazumder
arXiv AI
Jul 29

Bridging Compute- and Data-Optimal Pretraining

arXiv:2607. 25271v1 Announce Type: cross Abstract: Classical compute-optimal scaling laws assume an unbounded supply of fresh pretraining data, yet pretraining is increasingly entering a regime in which compute grows faster than the availability of high-quality data.

By Tian Qin, Kimia Hamidieh, David Alvarez-Melis
arXiv AI
3d ago

Revisiting scaling laws for reward optimization

The paper presents a new scaling law for reward optimization in AI alignment, showing that performance scales as Θ(√min{log(M), K}), where M is the number of preference comparisons used to train a proxy reward model and K is the KL‑divergence budget relative to a reference policy. The authors derive this law using an information‑theoretic model, prove its tightness, and validate it with extensive experiments involving a 70B gold reward model and smaller proxy models (0.6B–4B). The empirical results demonstrate a strong fit (R² 97–99 %) across different model sizes, noise levels, and optimization methods, suggesting that reward optimization behaves like a simple selection task over IID Gaussian variables with noisy feedback.

By Ali Aouad, Aymane El Gadarri, Vivek F. Farias
arXiv AI
Aug 25

Towards a Densing Law for User Representation Learning at Billion-Scale Capacity

The paper introduces a User Behavioral Densing Law that quantifies how the minimum sufficient tokenization capacity scales with data size in user representation learning. A pilot study on a billion‑scale Alipay dataset shows raw data scaling bottlenecks and the benefits of tokenization, while theoretical analysis and experiments reveal an approximately linear relationship between the logarithms of tokenization capacity and input data size. Using this law, the authors develop ALGN, an adaptive variable‑length tokenization method that outperforms existing baselines across diverse data sources and downstream tasks.

By Bin Dou, Junru Zhang, Zhaoyi Yuan, Wuliang Huang, Letian Gong, Baokun Wang, Huan Li, Yu Cheng, Weiqiang Wang
arXiv AI
Jun 8

Understanding Generative Recommendation with Semantic IDs from a Model-scaling View

arXiv:2509. 25522v3 Announce Type: replace Abstract: Recent advancements in generative models have allowed the emergence of a promising paradigm for recommender systems (RS), known as Generative Recommendation (GR), which tries to unify rich item semantics and collaborative filtering signals.

By Jingzhe Liu, Liam Collins, Jiliang Tang, Tong Zhao, Neil Shah, Clark Mingxuan Ju
arXiv AI
Sep 2

From Language to Behavior: Scaling Sequence Transformers for Industrial Recommendation Ranking with Rec-Native Designs

The paper introduces ReST, a recommendation‑native Transformer scaling framework designed to handle noisy, irregular, and sparsely supervised user behavior sequences in production ranking. ReST employs a dual‑gated attention encoder with rotary positional and temporal embeddings, and a lightweight cross decoder that decouples heavy encoding from fast decoding, enabling efficient compute‑once, decode‑many‑times ranking. Experiments on industrial and public benchmarks show that ReST outperforms traditional Transformer blocks, achieving higher accuracy and consistent scaling across sequence length, depth, and width, and a one‑week online A/B test on a production advertising platform yielded a 1.31% AUC lift and an 11.93% increase in a core revenue metric within a 50 ms P99 latency budget.

By Jie Chen, Xiangqian Yu, Yanchao Lian, Tan Lu, Run Yang, Zhengchun Shang, Xing Wang, Cheng Chen, Ke Hu, Qiang Li, Tianjiu Yin, Xiaobing Liu
arXiv AI
Jul 24

Benchmarking the Personalization Capabilities of Large Language Models

arXiv:2607. 20471v1 Announce Type: new Abstract: Personalization, the act of varying a message to induce action from a specific receiver while keeping sender, channel, and time fixed, has a long tradition in psychology and marketing as a two-party problem in which sender and receiver have independent objectives.

By Ashutosh Srivastava, Siddharth Yedlapati, Vinay Aggarwal, Yaman Kumar Singla, Shashwat Dixit, Jitendra Ajmera, Balaji Krishnamurthy
arXiv Machine Learning
Sep 14

Breaking the Token Ceiling: Distilling Smaller, Stronger Byte Models

The paper investigates whether small models distilled from larger ones behave similarly when using byte versus token tokenization. It introduces two methods—Marginalize‑It (approximate) and End‑Of‑Token (exact)—to convert token logits to byte logits, and conducts a large‑scale study on decoder‑only dense transformers ranging from 1 billion to 1 trillion bytes of data. Results show that while token‑based models excel early, byte‑based models eventually surpass them with more compute, achieving higher performance ceilings, greater data efficiency, and lower logit storage costs.

By Kalyani Marathe, Artidoro Pagnoni, Tomasz Limisiewicz, Margaret Li, Mike Lewis, Luke Zettlemoyer, Srinivasan Iyer