arXiv AI

Synthetic Data from Cross-Domain Events for Large-Scale Recommendation Systems

arXiv:2606. 00282v1 Announce Type: cross Abstract: Large-scale recommendation systems operate across diverse domains, yet they face the challenges of data sparsity and noisy implicit feedback.

arXiv Computation and Language
Sep 11

LLMAR: A Tuning-Free Recommendation Framework for Sparse and Text-Rich Industrial Domains

LLMAR is a tuning‑free recommendation framework designed for sparse, text‑rich industrial B2B domains. It transforms user behavioral history into structured semantic motives using LLM inference, employs a reflection loop to self‑correct hallucinations, and operates cost‑effectively with asynchronous batch processing. Experiments on MovieLens‑1M, Amazon Prime Pantry, and a construction risk dataset show LLMAR surpasses state‑of‑the‑art learning models, achieving up to a 54.6% nDCG@10 improvement while keeping inference costs around $1 per 1,000 users.

By Ryogo Hishikawa, Ichiro Kataoka, Shinya Yuda
arXiv AI
Jun 2

Principled Synthetic Data Enables the First Scaling Laws for LLMs in Recommendation

arXiv:2602. 07298v3 Announce Type: replace-cross Abstract: Large Language Models (LLMs) represent a promising frontier for recommender systems, yet their development has been impeded by the absence of predictable scaling laws, which are crucial for guiding research and optimizing resource allocation.

By Benyu Zhang, Qiang Zhang, Jianpeng Cheng, Hong-You Chen, Qifei Wang, Wei Sun, Shen Li, Jia Li, Jiahao Wu, Qunshu Zhang, Neeraj Bhatia, Xiangjun Fan, Hong Yan
arXiv AI
Sep 25

Cross-Country Code-Mixing for Generative Recommendation

Cross-Country Code-Mixing for Generative Recommendation (CMRec) is a framework that enhances generative recommendation across different countries by injecting cross-country supervision at the data level. It learns a shared semantic codebook from multi-modal content and behavioral co-occurrence, then synthesizes mixed-country sequences through token-level substitutions that respect both static and dynamic constraints. A context-aware loss reweights these mixed samples based on their plausibility, leading to improved recommendation quality in data-sparse countries while maintaining performance in data-rich markets, as demonstrated by significant gains in advertising revenue and orders in real-world e-commerce experiments.

By Yuan Gao, Hao Deng, Haibo Xing, Yi Xu, Lingyu Mu, Jinxin Hu, Yu Zhang, Xiaoyi Zeng
arXiv Computation and Language
Aug 25

A Scalable Cross-Domain Event Extraction System via a Unified Generative Training Framework

The paper introduces a scalable cross‑domain event extraction system built on a unified generative sequence‑to‑sequence framework. It jointly handles event detection and argument extraction, allowing both pipeline and end‑to‑end configurations. By fine‑tuning pretrained language models on multiple event datasets from diverse domains, the system retains domain‑specific semantics while generalizing across large, evolving label spaces, and offers a web‑based application for researchers to upload documents, extract events, visualize triggers and arguments, and compare configurations.

By Siting Liang, Omar Adjali, Omair Shahzad Bhatti, Daniel Sonntag