Personalized incentive allocation is vital for e-commerce, where uplift modeling is the standard for estimating Individual Treatment Effects (ITE). However, traditional models often fail in complex multi-seller environments with violations of the Stable Unit Treatment Value Assumption (SUTVA).
arXiv:2606. 26690v1 Announce Type: cross Abstract: In large-scale paid acquisition and growth advertising systems, production attribution outputs are widely used for daily budget allocation and channel diagnosis.
By Donghui Li, Bowen Yuan, Zili Yang, Qinxin Chen, Lijing Song
arXiv:2609.16407v1 Announce Type: cross
Abstract: On a delivery platform, personalized store ranking greatly influences what users find and order. Unlike digital-only domains, candidate stores are lo...
By Marcel Kurovski, Attila Nagy, Steffen Klempau, Aleksandr Fedintsev
arXiv:2602. 12972v2 Announce Type: replace-cross Abstract: In online advertising, marketing interventions such as coupons introduce significant confounding bias into Click-Through Rate (CTR) prediction.
By Siyun Yang, Shixiao Yang, Jian Wang, Di Fan, Kehe Cai, Haoyan Fu, Jiaming Zhang, Wenjin Wu, Peng Jiang
arXiv:2606. 30932v1 Announce Type: new Abstract: Two-sided marketplaces connect distinct user groups whose interests often conflict -- improving outcomes on one side could degrade the other side's experience.
By Yufei Wu, Zhen Yan
arXiv:2607. 28956v1 Announce Type: new Abstract: Large language model agents are increasingly evaluated as autonomous tool users, yet most benchmarks focus on bounded tasks with immediate success criteria.
By Qiming Shi, Yulong Tao, Linbo Jin, Zhaolu Kang, Yibo Dou, Jiawen Zhu, Tianjun Pan, Shaokang Fu, Chengyu Wang, Siyue Li, Yaping Cheng, Di Weng, Chengfu Huo
E-Commerce Bench is an open‑source benchmark that simulates a year‑long e‑commerce operation, requiring LLM agents to manage multiple online stores, negotiate with suppliers, optimize sales, fulfill orders, handle returns, and manage cash flow. The environment uses real product and supplier data, a calendar of promotions and shocks, and deterministic customer and negotiation models to enable reproducible evaluation. The study evaluates 18 state‑of‑the‑art models across seven metrics, finding no single model dominates, with GPT‑5.6 Sol achieving the highest year‑end assets but lagging in fraud avoidance and operational efficiency.
By Wei Fan, Xinjie Shen, Xudong Guo, Jianhong Tu, Yang Su, Yinger Zhang, Lianghao Deng, Fengyu Wang, Baohua Dong, Yangqiu Song, Dayiheng Liu
arXiv:2608. 11675v1 Announce Type: new Abstract: Coupon campaigns seek to lift both conversion and revenue, but gross merchandise value (GMV) follows a deterministic funnel from conversion to conditional order value and is zero-inflated and heavy-tailed.
By Yu Zhang (AMap Alibaba Group, Beijing, China), Zhihan Wang (AMap Alibaba Group, Beijing, China), Guanlin Chen (AMap Alibaba Group, Beijing, China), Min Jiang (AMap Alibaba Group, Beijing, China), Shuai Li (AMap Alibaba Group, Beijing, China)
CAVEAT is a new benchmark that tests computer‑use agents (CUAs) in nine online marketplace environments where platform incentives may steer agents away from user goals. The study finds that agents succeed in choosing user‑optimal products only 78.6% of the time in neutral settings, dropping to 17.3% when steering mechanisms are active. By diagnosing three failure points—priority distortion, premature narrowing of options, and early commitment—CAVEAT-Harness interventions raise user‑optimal purchasing success by 55.0%.
By Yuxuan Li, Will Epperson, Wesley Deng, Zezhou Huang
arXiv:2603. 20775v2 Announce Type: replace Abstract: In personalized marketing, uplift models estimate the incremental effect of an intervention by modeling how customer behavior would change under alternative treatments using counterfactual analysis.
By Yuxuan Yang, Dugang Liu, Yiyan Huang
The paper introduces DCEO, a data‑driven framework that learns item‑level proxy scores directly aligned with long‑term user objectives in e‑commerce search. It aggregates these scores into a user‑level metric, measures alignment via relative causal effect, and uses an actor‑critic model to generate context‑dependent fusion weights for multiple objectives. Offline experiments and a 41‑day online A/B test show DCEO improves GMV by 0.36% over traditional proxies.
By Junzhao Zhang, Tao Zhang, Liren Yu, Feiyi Dong, Zhixuan Zhang, Dan Ou, Haihong Tang
The paper introduces LLP, a Large Language Model–based generative framework for pricing second‑hand products on consumer‑to‑consumer platforms. LLP retrieves similar items to capture market dynamics, then uses LLMs to generate price suggestions, refined through supervised fine‑tuning and group relative policy optimization. A confidence‑based filter rejects unreliable predictions, and experiments show LLP outperforms prior methods, achieving higher static adoption rates when deployed on Xianyu.
By Hairu Wang, Sheng You, Qiheng Zhang, Xike Xie, Shuguang Han, Yuchen Wu, Fei Huang, Jufeng Chen