AgenticGen is a reward‑guided framework for generating advertising videos that splits the task into strategy selection and draft generation, allowing online business feedback to supervise each stage. It learns performance‑based rewards from accumulated online metrics and rubric‑based rewards aligned with human quality standards, then applies DPO and GRPO to refine policies. Offline tests confirm the reward models, and online A/B tests on TikTok show significant gains in CTR, CVR, and advertising value over a baseline.
By Xingyuan Bu, Chengru Song, Hao Zhou, Tao Zhou, Dong Li, Wei Li, Shilong Li, Hao Shi, Yongxin Guo, Donghao Zhou, Qiangpeng Yang, Shilei Wen
RL-ADA introduces a co‑evolutionary training framework that replaces costly human annotations with world‑feedback rewards derived from interaction outcomes. In this system, a large Customer Support Agent and an Adversarial Customer Agent train together, guided by an automated judge that rewards successful resolution and realistic intent‑concealing utterances, respectively. Applied to a banking support proof of concept, the method eliminates routing errors and doubles the end‑to‑end PASS rate over five cycles, while also revealing a new adversarial strategy called Contextual Camouflage.
By Ram Narayanan, Harshit Rajgarhia, Abhishek Mukherji
Dialogue systems in e-commerce scenarios often need to satisfy multiple objectives: accurately reasoning over user profiles (e. g.
arXiv:2607. 17281v1 Announce Type: cross Abstract: Auto-bidding plays an essential role in online advertising, automatically adjusting bids for advertisers to optimize their commercial goals.
By Yuejia Dou, Hesong Wang, Xinyu Zhang, Tianyu Wang, Zhilin Zhang, Chuan Yu, Jian Xu, Bo Zheng, Qi Qi
arXiv:2602. 10226v2 Announce Type: replace-cross Abstract: Optimizing large-scale machine learning systems, such as recommendation models for global video platforms, requires navigating a massive hyperparameter search space and, more critically, designing sophisticated optimizers, architectures, and reward functions to capture nuanced user behaviors.
By Haochen Wang, Yi Wu, Daryl Chang, Li Wei, Lukasz Heldt
arXiv:2601. 02871v3 Announce Type: replace Abstract: Task-oriented proactive dialogue agents play a pivotal role in recruitment, particularly for steering conversations towards specific business outcomes, such as acquiring social-media contacts for private-channel conversion.
By Zhiyong Cao, Dunqiang Liu, Qi Dai, Haojun Xu, Huai Yuen Khor, Hao Wang, Huan He, Yafei Liu, Ke Ma, Ruqian Shi, Sicheng Zhou, Sijia Yao
arXiv:2605. 28882v2 Announce Type: replace-cross Abstract: With the rapid advancement of large language models, evaluating human-likeness in open-ended conversation has become increasingly important.
By Yihang Lin, Yunze Gao, Zeyang Lin, Dongbo Li, Kun Peng, Yue Liu
The paper introduces a lifecycle framework for LLM-as-a-Judge systems used to evaluate recommendation explanations at Netflix. It outlines four phases—Birth, Training, Deployment, and Monitoring—detailing how each stage addresses specific technical and operational challenges. The authors report that after five weeks of A/B testing, judge-aligned explanations increased novel content viewing and successful browse-to-play sessions without quality takedowns.
By Emma Yanyang Kong, JJ Tan, Ishan Gupta, Lars Olds, Claire Campbell, David Fagnan, Veli Balin, Rohan Gosain, Louis Garcia, Minsu Jang
UniPolicy is a multi-policy alignment framework for search advertising that jointly optimizes relevance, click propensity, and commercial value. It uses objective-specific prefix tokens, sparse MoE-LoRA routing, and residual FFNs to decouple parameters within a shared backbone, and builds pairwise preferences from multi-stage behavioral feedback to improve generation. In large-scale offline tests and a 7‑day online A/B test, UniPolicy achieved balanced gains across metrics, boosting CTR by 0.71%, RPS by 1.58%, and revenue by 1.32% while keeping latency stable.
arXiv:2511. 19314v2 Announce Type: replace Abstract: Information-seeking is a core capability for AI agents, requiring them to gather and reason over tool-generated information across long trajectories.
By Jaewoo Lee, Archiki Prasad, Justin Chih-Yao Chen, Zaid Khan, Elias Stengel-Eskin, Mohit Bansal
arXiv:2608. 10209v1 Announce Type: new Abstract: Feedback signals used to train Large Language Models (LLMs) are the primary driver of their behavior and our main lever for instilling alignment with human values and objectives.
By Alec Harris, Kasey Corra, Archie Chaudhury, Yixiong Hao
The paper introduces SalesLLM, a bilingual (Chinese/English) benchmark for evaluating large language models (LLMs) in realistic sales dialogues. It comprises 30,074 scripted configurations and 1,805 curated multi‑turn scenarios from Financial Services and Consumer Goods, with controllable difficulty and personas. An automatic evaluation pipeline uses an LLM judge for sales‑process progress and fine‑tuned BERT classifiers for end‑of‑dialogue buying intent, while a user model, CustomerLM, is trained to improve simulation fidelity. SalesLLM scores correlate strongly with human ratings (Pearson r = 0.86) and reveal that top Chinese LLMs match junior‑to‑intermediate human salespeople but not experts, with cross‑lingual consistency remaining poor.
By Xuanbo Su, Wenhao Hu, Le Zhan, Yuting Xie, Kailin Lyu, Kaijie Chen, Ziwei Li, Yeqiang Wang, Haibo Su, Yunzhang Chen, Ling Huang