arXiv AI

A Better Spur Should Start From Each Objective

arXiv Computation and Language
2d ago

UniPolicy: Unified Objective-Specific Policies for Generative Search Advertising

UniPolicy is a unified objective‑specific policy framework for search advertising that jointly optimizes relevance, click propensity, and commercial value. It uses objective‑aware prefix tokens, sparse MoE‑LoRA routing, and residual FFNs to decouple parameters within a shared backbone, and constructs pairwise preferences from multi‑stage behavioral feedback to strengthen clicked candidates. In large‑scale offline tests and a 7‑day online A/B test, UniPolicy improves CTR by 0.71%, RPS by 1.58%, and advertising revenue by 1.32% while keeping serving latency stable.

By Kun Yao, Yuhang Zhou, Yichi Zhang, Zeliang Tong, Shengri Xue, Haitao Wang, Siyu Lu, Qianlong Xie, Xingxing Wang
arXiv Machine Learning
Sep 3

Objective-Behavior Alignment: Diagnostics for MORL Policy Selection

The paper introduces a diagnostic workflow for multi‑objective reinforcement learning (MORL) that reveals behavioral differences among policies on the Pareto front, which are not apparent from value vectors alone. It offers quantitative and visual tools to inspect these variations and demonstrates their effectiveness on both simple grid tasks and more complex continuous‑control benchmarks.

By Antonio Mone, Zuzanna Osika, Florian Felten, Pradeep K. Murukannaiah, Mark Fuge, Frans A. Oliehoek, Luciano Cavalcante Siebert
arXiv Computation and Language
5d ago

Optimizing Sparse Outcomes Through Dense Behavioral Signals via Value-Guided Preference Distillation

arXiv:2609.14648v1 Announce Type: new Abstract: Aligning multi-turn dialogue agents is usually framed as matching turn-level human preferences, yet direct optimization of long-term outcomes is often...

By Ziyi Zhu, Daniel R. Cahn, Thomas D. Hull, Caitlin A. Stamatis, Olivier Tieleman, Guilherme B. Freire, Jinghong Chen