Offline-to-Online Creative Optimization with Generative Models and Adaptive Testing
arXiv:2607. 23696v1 Announce Type: new Abstract: Ad creative optimization is increasingly constrained by evaluation rather than generation.
PinDCO is a scalable dynamic creative optimization system designed for Pinterest’s billion‑scale visual discovery platform. It uses a Creative Component Fusion Network to score ad creatives by modeling individual components (image, title, layout) with dedicated towers and fusing their representations, while a Pixel‑aware Adjustment Module tailors scores to creative size for better whole‑page outcomes. The system incorporates a lightweight pre‑selection model, caching, and dynamic batching to handle large candidate volumes, achieving a 3.09% lift in ad click‑through rate in online experiments.
arXiv:2607. 23696v1 Announce Type: new Abstract: Ad creative optimization is increasingly constrained by evaluation rather than generation.
arXiv:2603. 11863v2 Announce Type: replace Abstract: The saturation of high-quality pre-training data has shifted research focus toward evolutionary systems capable of continuously generating novel artifacts, leading to the success of AlphaEvolve.
arXiv:2606. 31711v1 Announce Type: new Abstract: Faithfulness -- how precisely a generated image aligns with its prompt -- is increasingly central to the real-world utility of text-to-image (T2I) models.
arXiv:2607. 20528v1 Announce Type: new Abstract: Online recommendation platforms increasingly use Large Language Models (LLMs) to extract structured features from ad creatives.
CommerceVibe is a system that generates e‑commerce creatives by synthesizing executable HTML/CSS code conditioned on product images, design requirements, and product information. It uses dual‑feedback reinforcement learning, combining rule‑based checks for text readability, product visibility, and layout validity with visual feedback from a vision‑language model that evaluates perceptual and commercial aspects. After fine‑tuning a large language model on 28,000 examples and applying dual‑feedback reinforcement learning, CommerceVibe achieves a weighted score of 94.0/100 on a 1,300‑case benchmark, outperforming both its SFT‑only counterpart and external models, and is validated by expert blind evaluations.
Large language models (LLMs) have achieved remarkable progress in language understanding, reasoning, and generation, sparking growing interest in their creative potential. Realizing this potential requires systematic and scalable methods for evaluating creativity across diverse tasks.
arXiv:2606. 11762v1 Announce Type: cross Abstract: Large language models (LLMs) have achieved remarkable progress in language understanding, reasoning, and generation, sparking growing interest in their creative potential.
DeGRe is a dense‑supervised generative reranking framework designed to improve multi‑stage recommender systems by addressing label bias and credit assignment issues. It uses an offline Lookahead Evaluator with beam search to generate dense supervision signals, which are distilled into a lightweight Online Generator that can perform efficient greedy decoding at inference time. Experiments show that DeGRe outperforms baselines on public benchmarks and industrial datasets, and it has been successfully deployed on Taobao Flash Shopping to enhance online recommendations.
arXiv:2607. 02484v1 Announce Type: cross Abstract: Visual token pruning is a crucial strategy for accelerating VLMs by compressing redundant image patches, yet existing methods often fail to preserve critical cues under dense instructions and fine-grained queries.
KwaiMind is a commercial image editing system that combines general editing capabilities with e-commerce specialization. It uses an agent-based data engine with 1.8 million editing pairs and a multimodal diffusion transformer trained through pre‑training, fine‑tuning, preference optimization, and online reinforcement learning. The system is guided by a vision‑language judge and specialized rewards for click‑through rate, text rendering, and product consistency, and it achieves top scores on ImgEdit, GEdit, REDEdit, and the new Ecom‑Bench, while improving predicted and actual CTR in offline and online experiments.
Editable Visual Design introduces a new design paradigm that combines a Coding Agent with a Vision‑Language Model (VLM) and an image generation model. The VLM acts as the creative brain, understanding requirements, planning tasks, and judging aesthetics, while the image generator produces isolated visual assets on demand. The agent follows an "imagine first, then act" workflow, generating assets, writing native HTML/CSS, and refining the design through visual feedback, ultimately producing editable, layer‑wise artifacts with real text that can be adjusted via a graphical interface.
arXiv:2608. 07243v1 Announce Type: new Abstract: Generative models are often evaluated through singular artifacts, whereas human creativity typically emerges through iterative generation, appraisal, and refinement.