arXiv Machine Learning

BRiG-AFA: Bellman Risk-to-Go Learning for Non-Myopic Active Feature Acquisition

arXiv:2608. 02305v1 Announce Type: new Abstract: Active feature acquisition (AFA) asks which unobserved feature to measure next for each test instance under a budget.

arXiv Machine Learning
Sep 10

Large Classification-Risk-Optional Label Acquisition

arXiv:2609.06873v1 Announce Type: cross Abstract: We study how a limited labeling budget should be allocated to minimize multiclass zero-one classification risk. We consider parametric classification...

By F. Setoudehtanzangi, Geoffrey J. McLachlan
arXiv Machine Learning
Sep 15

Learning Source Acquisition Policies by Offline Planning

The paper introduces O-MPAC, an offline planning method for learning source acquisition policies under a limited budget. It transfers finite‑horizon risk‑cost targets from full training data into a shared source‑action scorer that re‑evaluates partial observations and source metadata after each query, applying a hard cost mask. Experiments show that O‑MPAC achieves high accuracy (0.965) in a routing task and outperforms several baselines on six real tasks, achieving the highest mean budget‑integrated accuracy on five of them.

By Ziqi Zhao, Run Xu, Qingjian Ni
arXiv Machine Learning
Aug 13

When Offline Evaluation Misleads: A Diagnostic Protocol for Reward and Policy Selection in Delayed-Feedback Contextual Bandits

arXiv:2608. 11560v1 Announce Type: new Abstract: Personalizing marketing messages with contextual multi-armed bandits (CMABs) drives real business value, yet the objective that ultimately matters - a downstream conversion - is observed only weeks later, too late to drive online learning.

By Sang Su Lee, Vineeth Loganathan, Shishir Dash, Vijay Raghavan
arXiv Machine Learning
Aug 31

Budget-Constrained Causal Bandits: Bridging Uplift Modeling and Sequential Decision-Making

The paper introduces Budget-Constrained Causal Bandits (BCCB), an online framework that learns individual treatment effects, explores uncertain users, and manages budget pacing simultaneously. It derives a per-arrival decision rule from a KKT condition of a Lagrangian relaxation, providing a principled algorithmic foundation. Experiments on the Criteo Uplift dataset show BCCB outperforms offline pipelines and other online baselines, especially when historical data is scarce (below 7,500 observations).

By Abhirami Pillai
Hugging Face Trending Papers
Aug 12

When Offline Evaluation Misleads: A Diagnostic Protocol for Reward and Policy Selection in Delayed-Feedback Contextual Bandits

Personalizing marketing messages with contextual multi-armed bandits (CMABs) drives real business value, yet the objective that ultimately matters - a downstream conversion - is observed only weeks later, too late to drive online learning. Teams therefore train the bandit on a fast proxy reward, and separately must judge whether a contextual bandit is worth its complexity over sending one best message.

arXiv Machine Learning
Jun 3

Human-in-the-Loop Contextual Bandits for Short-Term Rental Dynamic Pricing: Structural Equivalence of Historical Warm-Up and Approval-Gated Live Learning

arXiv:2606. 02595v1 Announce Type: new Abstract: Dynamic pricing in short-term rental (STR) markets presents a distinctive challenge for online learning algorithms: pricing decisions carry significant financial risk, operators require explainability, and market feedback is sparse (one booking outcome per listed night).

By Oleg Miroshnichenko