Incentivized advertising allocates monetary or virtual rewards to drive user engagement, where a key challenge is optimizing continuous incentive magnitudes under strict global constraints. This problem is complicated by high-frequency interactions, delayed feedback, and non-Markovian user dynamics such as fatigue, which limit the effectiveness of existing uplift modeling and constrained reinforcement learning approaches.
arXiv:2602. 08261v2 Announce Type: replace Abstract: Auto-bidding systems strive to maximize marketing value while maintaining high compliance with efficiency constraints, such as Target Cost-Per-Action (CPA).
By Binglin Wu, Yingyi Zhang, Xianneng Li, Ruyue Deng, Chuan Yue, Weiru Zhang, Xiaoyi Zeng
arXiv:2608. 10182v1 Announce Type: cross Abstract: Large-scale targeting and recommendation systems are typically built around predictive scores fed into heuristic or local allocation.
By Changshuai Wei, John Bencina, Phuc Nguyen, Andre Assuncao Silva T Ribeiro, Benjamin Zelditch
arXiv:2608. 19842v1 Announce Type: new Abstract: Agentic reinforcement learning (RL) has become a critical stage in the post-training of large language models.
By Dayang Liang, Lang Feng, Bo An, Yunlong Liu
FLARE introduces a dense supervision paradigm for long‑horizon coding agents, leveraging a Generative Reward Model (GRM) trained via the RADAR diagnostic framework. The GRM provides real‑time, step‑level risk feedback, enabling FLARE to act as an active scaffold that intercepts high‑risk steps during inference and supplies structured signals for post‑training fine‑tuning and reinforcement learning. Experiments show FLARE outperforms existing methods, achieving a 5× reduction in token consumption and significant performance gains in both supervised fine‑tuning and RL settings.
By Jingxuan Xu, Gang Wu, Yanan Wu, Yutao Mou, Songwei Yu, Tianzhuang He, Zhengshuo Gong, Zhao Liu, Zihang Xu, Wenqiang Zhu, Xinping Lei, Weihao Li, Yuhui Bai, Zhongqiu Wang, Yan Wu, Ariel Deng
OneBid is a unified auto‑bidding foundation model that consolidates diverse cost‑per‑X (oCPX) advertising scenarios into a single framework. It builds on Decision Transformer by conditioning on two atomic signals—Return‑to‑Go for conversion value and Cost‑to‑Go for cost ratio—and incorporates value‑aware regularization. A sequence‑level Mixture‑of‑Experts architecture captures cross‑scenario knowledge while preserving low latency, and a Critic‑guided Relative Offline Policy optimization (CROP) aligns the backbone with scenario‑specific preferences without unsafe online exploration. In production at Kuaishou, OneBid achieved a 2.2% overall ADVV increase and up to 13.1% in the ROAS scenario.
By Yewen Li, Peng Jiang, Yitian Li, Pengfei Lv, Xialong Liu, Peng Jiang, Qingpeng Cai