arXiv AI

Ad Headline Generation using Self-Critical Masked Language Model

arXiv:2607. 06818v1 Announce Type: cross Abstract: For any E-commerce website it is a nontrivial problem to build enduring advertisements that attract shoppers.

arXiv Machine Learning
Aug 28

Token-Level Advertising

The paper introduces the Latent Advertiser Mixture Auction (LAMA), a token‑level advertising framework that integrates advertiser influence directly into the text generation process. Advertisers provide local continuation values that shape next‑token policies, and the platform decodes these through a latent mixture while updating an allocation posterior. LAMA is shown to satisfy Markov DSIC and IR, achieve near‑optimal KL‑regularized welfare, and, in proof‑of‑concept experiments on commercial‑search queries, improve platform welfare and revenue without compromising user‑facing response quality.

By Hanbing Liu, Bowei Zhang, Changyuan Yu, Yinyu Ye, Qi Qi
arXiv Computation and Language
Sep 10

AgenticGen: Reward-Guided Agentic Video Generation for Advertising

AgenticGen is a reward‑guided framework for generating advertising videos that splits the task into strategy selection and draft generation, allowing online business feedback to supervise each stage. It learns performance‑based rewards from accumulated online metrics and rubric‑based rewards aligned with human quality standards, then applies DPO and GRPO to refine policies. Offline tests confirm the reward models, and online A/B tests on TikTok show significant gains in CTR, CVR, and advertising value over a baseline.

By Xingyuan Bu, Chengru Song, Hao Zhou, Tao Zhou, Dong Li, Wei Li, Shilong Li, Hao Shi, Yongxin Guo, Donghao Zhou, Qiangpeng Yang, Shilei Wen
arXiv Machine Learning
Sep 3

DMRL: Document-Mediated Reinforcement Learning for Skill Optimization in Advertising Recommendation

The paper introduces Document-Mediated Reinforcement Learning (DMRL), a framework that treats skill document optimization in advertising recommendation as a sequence of structured editing actions. An upper-level agent edits the skill documents while a frozen lower-level task agent evaluates the edits via A/B testing. DMRL incorporates Dual-Relative Policy Optimization for robust advantage estimation and a Long-term Reward Predictor that models population heterogeneity to estimate long-term outcomes, achieving superior performance on a large-scale short‑video ads platform.

By Wei Zhang, Hongji Li, Song Sun, Peng Yu, Xue Yang, Lei Zhao, Peng Jiang
arXiv AI
Sep 1

Agent2UCB: Agentic System for Generative Engine Optimization

Agent2UCB is a new agentic system designed for Generative Engine Optimization (GEO), which refines content to boost its likelihood of being cited or summarized by generative AI search engines. The system autonomously evaluates nine GEO strategies for each content item, selects the most effective one, and speeds up this selection using a bandit-based Agent2UCB policy that blends large language model priors with real-time reward signals. Additionally, it offers a lightweight, text-only SEO readiness check that assesses readability, topical coverage, and EEAT-style credibility, and experiments on GEO-Bench demonstrate consistent visibility gains while maintaining SEO quality.

By Sheldon Yu, Rui Wang, Tong Yu, Sungchul Kim, Doga Dogan, Junda Wu, Julian McAuley
arXiv Machine Learning
Jun 15

DRIVE: Distributional and Retrieval-Augmented Bidding with Value Evaluation

arXiv:2606. 14192v1 Announce Type: new Abstract: Auto-bidding is a core component of real-time advertising systems, where decisions must optimize long-term performance under budget and cost constraints, while online exploration is prohibitively risky.

By Miduo Cui, Haochen Wang, Shangqin Mao, Xun Yang, Qianlong Xie, Xingxing Wang, Xuri Ge, Ying Zhou, Zhiwei Xu