arXiv AI

Selection Bias Correction in Retail Intelligence

The paper examines how retail intelligence, which often focuses on high‑velocity products, can suffer from selection bias that skews inflation estimates by overlooking niche items. Using 400 Monte Carlo simulations across four data‑generating scenarios, the authors compare Inverse Probability Weighting (IPW) and stratification methods. They find that stratification generally outperforms IPW—achieving sub‑0.04 percentage‑point median error even when population breaks misalign—while IPW only excels under smooth polynomial relationships, highlighting the importance of method choice in long‑tail retail contexts.

arXiv Machine Learning
Jul 30

Lottery Tickets Are Not Deployment Tickets

arXiv:2607. 27031v1 Announce Type: new Abstract: Reports on how sparsification, compression, and lottery tickets change model behavior have been mixed in the prior literature, with beneficial effects observed in some studies and adverse effects in others.

By Bum Jun Kim
arXiv AI
Jun 2

Generative AI and Sales Productivity: Field Experiments in Online Retail

arXiv:2510. 12049v4 Announce Type: replace-cross Abstract: We quantify the short-term impact of Generative Artificial Intelligence (GenAI) on sales performance through a series of large-scale randomized field experiments involving millions of users and products at a leading cross-border online retail platform.

By Lu Fang, Zhe Yuan, Kaifu Zhang, Dante Donati, Miklos Sarvary
arXiv Machine Learning
Aug 31

Robust Assortment Optimization from Observational Data

The paper introduces a robust framework for assortment optimization that addresses distributional shifts in customer choice behavior. It demonstrates computational tractability when the nominal choice model is known and develops statistically optimal algorithms for the data‑driven setting, providing matching upper and lower bounds on sample complexity. The authors identify "robust item‑wise coverage" as the minimal data requirement for efficient robust learning, bridging robustness and statistical efficiency in assortment planning.

By Miao Lu, Yuxuan Han, Han Zhong, Zhengyuan Zhou, Jose Blanchet
arXiv Machine Learning
Aug 31

Timing-Aware Repurchase Prediction for Web-Scale E-Commerce: Survival Models for Multi-Surface Grocery Recommendation

The paper proposes replacing multiple horizon‑specific binary classifiers with a single survival model to predict time‑to‑repurchase in grocery e‑commerce. Empirical analysis shows a slightly decreasing hazard (k≈0.9) and that a Log‑Normal model best fits marginal distributions while Weibull best fits residuals. A single Accelerated Failure Time (AFT) model matches or surpasses per‑horizon classifiers with fewer trees, and a 4‑parameter calibration maps survival CDFs to horizon probabilities without monotonicity violations, revealing a trade‑off between calibration and ranking within the AFT family.

By Akshay Kekuda, Shreeranjani Srirangamsridharan, Ishan Bhatt, Yanan Cao, Sinduja Subramaniam, Evren Korpeoglu, Kaushiki Nag, Kannan Achan