arXiv AI

Beyond Interaction Capacity: Estimator Scaling with Recursive Models for CTR Prediction

arXiv AI
Jun 8

Design Once, Deploy at Scale: Template-Driven ML Development for Large Model Ecosystems

arXiv:2603. 24963v3 Announce Type: replace Abstract: Modern computational advertising platforms typically rely on recommendation systems to predict user responses, such as click-through rates, conversion rates, and other optimization events.

By Jiang Liu, John Martabano Landy, Yao Xuan, Swamy Muddu, Nhat Le, Munaf Sahaf, Luc Kien Hang, Rupinder Khandpour, Kevin De Angeli, Chang Yang, Shouyuan Chen, Shiblee Sadik, Anirudh Agrawal, Djordje Gligorijevic, Jingzheng Qin, Peggy Yao, Alireza Vahdatpour
arXiv AI
Sep 25

DeGRe: Dense-supervised Generative Reranking for Recommendation

DeGRe is a dense‑supervised generative reranking framework designed to improve multi‑stage recommender systems by addressing label bias and credit assignment issues. It uses an offline Lookahead Evaluator with beam search to generate dense supervision signals, which are distilled into a lightweight Online Generator that can perform efficient greedy decoding at inference time. Experiments show that DeGRe outperforms baselines on public benchmarks and industrial datasets, and it has been successfully deployed on Taobao Flash Shopping to enhance online recommendations.

By Chaotian Song, Jingyao Zhang, Chenghao Chen, Zisen Sang, Dehai Zhao, Guodong Cao, Boxi Wu, Deng Cai, Jia Jia
arXiv Machine Learning
Jul 31

ROCS: Request-Oriented Compute Sharing for Efficient Large-Scale Recommendation

arXiv:2607. 27744v1 Announce Type: new Abstract: Modern recommendation models gain prediction quality by scaling feature-interaction and sequence modules, but production cost constraints cap how far systems can scale.

By Yuxin Chen, Liang Luo, Buyun Zhang, Jian Jiao, Boda Li, Haoyu Wang, Tongyi Tang, Ao Cai, Zijian Shen, Zhengkai Zhang, Wenyi Xie, Ryan Dick, Han Liu, Neng Shi, Bin Yu, Jianbo Xiao, Shuyao Bi, Hongtao Yu, Yuanwei Fang, Zhuoran Zhao, Sijia Chen, Yang Chen, Shuqi Yang, Qianru Li, Zikun Liu, Wei Ling, Sihan Zeng, Longhao Jin, Jiaxin Lu, Yinbin Ma, Jiawei Li, Yichen Ruan, Yong Ler Lee, Birmingham Guan, Zijian Li, Jianbo Sun, Zhengyu Zhang, Zeliang Chen, Xiaohan Wei, Yuchen Hao, GP Musumeci, Venkatesh Ranganathan, Yantao Yao, Chunqiang Tang, Wenlin Chen, Santanu Kolay, Ellie Dingqiao Wen
arXiv AI
Sep 10

MoEMB: Scaling Universal Multimodal Embeddings with Efficient Mixture-of-Experts Models

MoEMB introduces a mixture‑of‑experts (MoE) approach to scale universal multimodal embeddings (UME) without increasing the size of the output vector or relying on autoregressive decoding. By expanding encoder capacity along the expert axis, MoEMB achieves state‑of‑the‑art performance on MMEB‑V2 and MRMR benchmarks with only 3 B active parameters, outperforming TTE‑based methods that use more than four times as many active parameters and require significantly more compute. The paper also presents the first comprehensive study of adaptive computation for MoE‑based embeddings, exploring training‑time and inference‑time strategies to further improve efficiency for large‑scale retrieval and recommendation systems.

By Xuanming Cui, Shlok Kumar Mishra, Wentao Bao, Aashu Singh, Zihao Wang, Xiangjun Fan, Jun Xiao, Ser-Nam Lim, Jianpeng Cheng