arXiv Machine Learning

Attributed, But Not Incremental: Cannibalization-Corrected Attribution for Large-Scale Advertising

arXiv:2606. 26690v1 Announce Type: cross Abstract: In large-scale paid acquisition and growth advertising systems, production attribution outputs are widely used for daily budget allocation and channel diagnosis.

arXiv Machine Learning
Aug 12

MARCO: Click-Intent Decomposition for Calibrated Ads Conversion Prediction

arXiv:2608. 10562v1 Announce Type: new Abstract: Not all clicks are equal.

By Shiwen Shen, Xiru Huang, Liang Luo, Jianbo Sun, He Lyu, Zihang Fu, Ivonne Xu, Zhizhuo Li, Zhengyu Zhang, Pei-Ju Sung, Yunmiao Wang, Zixuan Wang, Zhengli Zhao, Qiang Jin, Mike Jermann, Mingda Li, Yang Xiao, Bhavana Challa, Brooke Bian, Yang Li, Ashish Chamoli, Bibek Bhusal, Danning Di, Yuan Jin, Meet Raval, Zhiwen Chen, Boyao Sun, Shuguang Wang, Yunlong He, Yantao Yao, Sagar Chordia, Wenlin Chen, Santanu Kolay, Qin Huang, Ellie Wen
arXiv AI
Sep 23

Testing, not presuming, adequacy: calibrating generative social simulators against emergent network structure

The paper introduces an adequacy‑aware calibration protocol for generative social simulators that integrates amortized posterior estimation, synthetic identifiability assessment, matched‑sample‑size adequacy checks, diagnosis‑guided repair, and held‑out audits. Applied to a second‑hand luxury resale market, the protocol reveals that behavioural parameters are recoverable but calibration is approximate and overconfident for one parameter, and that the simulator’s reachability reference is violated in every cell, particularly in mean purchased tier. The repair improves two of four cells but fails to restore full adequacy, and a held‑out audit uncovers a buyer‑breadth‑dispersion miss not detected earlier; profile‑source ablation shows language‑model‑derived persona profiles outperform a flat rule baseline, though within‑category brand relabelling has no consistent effect. whyItMatters":"The study demonstrates that without an adequacy check, generative social models may appear valid descriptively yet fail to capture key emergent network structures, highlighting the need for rigorous calibration protocols in social simulation research."

By Tengfei Shao, Chao Li, Xu Wang, Masayuki Goto
arXiv AI
Aug 28

Selection Bias Correction in Retail Intelligence

The paper examines how retail intelligence, which often focuses on high‑velocity products, can suffer from selection bias that skews inflation estimates by overlooking niche items. Using 400 Monte Carlo simulations across four data‑generating scenarios, the authors compare Inverse Probability Weighting (IPW) and stratification methods. They find that stratification generally outperforms IPW—achieving sub‑0.04 percentage‑point median error even when population breaks misalign—while IPW only excels under smooth polynomial relationships, highlighting the importance of method choice in long‑tail retail contexts.

By Spandan Ghose Chowdhury
arXiv Machine Learning
Sep 4

From Leakage to Fidelity: Reliable Benchmarking for Temporal Cascade Prediction

The paper critiques current temporal cascade prediction benchmarks for relying on leakage-prone random splits, limited datasets, and unexamined protocol effects. It proposes a fidelity-aware benchmarking suite featuring the Full Temporal protocol, overlap-based leakage diagnostics, and analyses of performance inflation and temporal drift. Additionally, it introduces the Taoke e‑commerce dataset with rich features and purchase conversions, and presents CasTemp as a lightweight reference method for scalable evaluation.

By Jie Peng, Rui Wang, Qiang Wang, Zhewei Wei, Bin Tong, Guan Wang, Bo Zheng