A Better Spur Should Start From Each Objective
Read the original on arXiv AI →The Flow has not summarised this story yet — read it at arXiv AI.
The Flow has not summarised this story yet — read it at arXiv AI.
arXiv:2505. 10892v2 Announce Type: replace Abstract: Post-training LLMs with RLHF and preference optimization methods (e.
arXiv:2607. 29559v1 Announce Type: new Abstract: Reinforcement Learning (RL) systems are typically trained using a single, well-specified scalar reward function.
arXiv:2602. 07764v2 Announce Type: replace-cross Abstract: Multi-objective reinforcement learning (MORL) seeks to train agents capable of balancing conflicting objectives.
UniPolicy is a unified objective‑specific policy framework for search advertising that jointly optimizes relevance, click propensity, and commercial value. It uses objective‑aware prefix tokens, sparse MoE‑LoRA routing, and residual FFNs to decouple parameters within a shared backbone, and constructs pairwise preferences from multi‑stage behavioral feedback to strengthen clicked candidates. In large‑scale offline tests and a 7‑day online A/B test, UniPolicy improves CTR by 0.71%, RPS by 1.58%, and advertising revenue by 1.32% while keeping serving latency stable.
arXiv:2506. 13702v4 Announce Type: replace-cross Abstract: Single-trajectory preference optimization methods learn from datasets of ((prompt, response, reward)) tuples, offering a practical alternative to pairwise preference learning by directly leveraging scalar feedback.
Search advertising connects user intent with commercial content and plays a critical role in platform monetization. Recent systems typically align pretrained generative models with a single business r...