arXiv Machine Learning

Robust Budget Pacing with a Single Sample

arXiv Machine Learning
Aug 31

Budget-Constrained Causal Bandits: Bridging Uplift Modeling and Sequential Decision-Making

The paper introduces Budget-Constrained Causal Bandits (BCCB), an online framework that learns individual treatment effects, explores uncertain users, and manages budget pacing simultaneously. It derives a per-arrival decision rule from a KKT condition of a Lagrangian relaxation, providing a principled algorithmic foundation. Experiments on the Criteo Uplift dataset show BCCB outperforms offline pipelines and other online baselines, especially when historical data is scarce (below 7,500 observations).

By Abhirami Pillai
arXiv Machine Learning
Aug 18

Dynamic Pricing and Advertising with Demand Learning

arXiv:2304. 14385v4 Announce Type: replace-cross Abstract: We consider a novel pricing and advertising framework in which a seller not only sets the product price but also designs flexible advertising schemes to influence customers' valuations of the product.

By Shipra Agrawal, Yiding Feng, Wei Tang
arXiv Machine Learning
22h ago

Online Generalized-Mean Welfare Maximization: Achieving Near-Optimal Regret from Samples

The paper investigates online fair allocation of sequential items to agents with heterogeneous preferences, aiming to maximize generalized-mean welfare. In an i.i.d. arrival setting, a pure greedy algorithm achieves near-optimal “~O(1/T)” average regret without needing distributional knowledge. For nonstationary arrivals, the authors show that a single historical sample per distribution suffices to recover the same regret rate, using re-solving algorithms that remain robust to distribution shifts.

By Zongjun Yang, Rachitesh Kumar, Christian Kroer
arXiv Machine Learning
Aug 3

Parameter-Free Heavy-Tailed Bandits

arXiv:2607. 29460v1 Announce Type: new Abstract: Heavy-tailed distributions arise naturally in sequential decision-making problems such as financial investment, online advertising, and network management, where rare but extreme outcomes can dominate performance.

By Gianmarco Genalti, Alberto Maria Metelli
arXiv Machine Learning
Sep 22

Leveraging Inference-Time Compute for Diffusion Models via Global Scheduling of Denoising Trajectories

The paper studies how to allocate a fixed computational budget across the denoising steps of diffusion models to improve sample quality at deployment. It shows that the expected benefit of evaluating multiple candidates at a step can be decomposed into a step‑specific sensitivity and a universal sample‑size factor, and that the optimal allocation follows a water‑filling structure. Experiments demonstrate that this allocation achieves the same quality as a uniform strategy while reducing function evaluations by 20–50%.

By Yuan Cao, Yifu Tang, Hangqi Li, Zeyu Zheng