Learning Variable-Length Tokenization for Generative Recommendation
arXiv:2605. 17779v2 Announce Type: replace Abstract: Generative recommendation reformulates recommendation as next-token prediction over discrete semantic identifiers (IDs).
The paper introduces RQ-Reg, a residual‑quantization framework for predicting continuous values in recommender systems. It decomposes target values into a sequence of quantization codes, autoregressively refining predictions from coarse to fine granularity, and incorporates an ordinal‑aware objective to align embeddings with target order. Experiments on watch‑time, LTV, and a large‑scale online A/B test for GMV demonstrate competitive performance and strong generalization across diverse prediction tasks.
arXiv:2605. 17779v2 Announce Type: replace Abstract: Generative recommendation reformulates recommendation as next-token prediction over discrete semantic identifiers (IDs).
arXiv:2604. 25834v2 Announce Type: replace Abstract: With the rapid development of the Internet, users have increasingly higher expectations for the recommendation accuracy of online content consumption platforms.
arXiv:2609.13789v1 Announce Type: new Abstract: In multi-channel paid user acquisition, early and accurate prediction of user retention at the channel level is crucial for optimizing budget allocatio...
SequenceO1 is an end‑to‑end framework that enables ultra‑long (up to 100K interactions) sequence modeling for recommendation systems. It compresses raw user histories into a fixed‑size sketch using Sketch Attention and then models short‑term and long‑term interests with Target‑to‑History Cross Attention. The system incorporates low‑rank caching, batching, pipeline lift, and a FlashSA kernel to keep training and inference efficient, achieving consistent offline and online performance gains when deployed at full traffic on Douyin.
ProtoFlow is a new multivariate time series forecasting framework that combines vector‑quantized autoencoding with prototype‑guided flow matching. It maps sequences into a discrete latent space, constructs a structured prior from the learned VQ codebook, and trains a DiT‑based rectified flow to transport samples from this prior to future latent representations conditioned on past observations. By replacing generic Gaussian noise with a learned prototype prior, ProtoFlow eliminates autoregressive rollout mismatch and achieves faster training convergence while delivering superior forecasting performance on benchmark datasets.
arXiv:2607. 17017v1 Announce Type: cross Abstract: As scalability becomes increasingly important in recommendation modeling, recent architectures have advanced the modeling of two broad sources of ranking signals along separate paths: non-sequence features, including user, item, context, and cross features; and sequence features from user behavior histories.
arXiv:2512. 10388v3 Announce Type: replace-cross Abstract: Conventional Sequential Recommender Systems (SRS) typically assign unique hash IDs (HID) to construct item embeddings, which mainly capture collaborative signals from historical user-item interactions.
arXiv:2609.35783v1 Announce Type: cross Abstract: Large-scale recommender systems, particularly short-form video platforms, are often bottlenecked by massive popularity feedback loops. In such enviro...
arXiv:2606. 01352v1 Announce Type: new Abstract: Watch time has emerged as a pivotal metric for optimizing deep user engagement in short-video recommender systems.
The paper introduces ReST, a recommendation‑native Transformer scaling framework designed to handle noisy, irregular, and sparsely supervised user behavior sequences in production ranking. ReST employs a dual‑gated attention encoder with rotary positional and temporal embeddings, and a lightweight cross decoder that decouples heavy encoding from fast decoding, enabling efficient compute‑once, decode‑many‑times ranking. Experiments on industrial and public benchmarks show that ReST outperforms traditional Transformer blocks, achieving higher accuracy and consistent scaling across sequence length, depth, and width, and a one‑week online A/B test on a production advertising platform yielded a 1.31% AUC lift and an 11.93% increase in a core revenue metric within a 50 ms P99 latency budget.
arXiv:2607. 25209v1 Announce Type: cross Abstract: Generative recommendation commonly represents items using fixed-length semantic identifiers (SIDs) constructed through clustering and quantization.
arXiv:2606. 07546v1 Announce Type: cross Abstract: Capturing user interests across extensive watch histories is critical for short-form video recommendation, yet scaling sequence length is limited by two bottlenecks: the semantic sparsity of atomic Video IDs and the quadratic computational complexity of Transformers.