The Challenger: When Do New Data Sources Justify Switching Machine Learning Models?
arXiv:2512. 18390v2 Announce Type: replace Abstract: Organizations often have an incumbent predictive model in production when new data sources become available.
arXiv:2608. 02305v1 Announce Type: new Abstract: Active feature acquisition (AFA) asks which unobserved feature to measure next for each test instance under a budget.
arXiv:2512. 18390v2 Announce Type: replace Abstract: Organizations often have an incumbent predictive model in production when new data sources become available.
arXiv:2609.06873v1 Announce Type: cross Abstract: We study how a limited labeling budget should be allocated to minimize multiclass zero-one classification risk. We consider parametric classification...
arXiv:2602. 05459v2 Announce Type: replace Abstract: Offline goal-conditioned reinforcement learning (GCRL) is typically benchmarked by the best tuned success rate of each method.
The paper introduces O-MPAC, an offline planning method for learning source acquisition policies under a limited budget. It transfers finite‑horizon risk‑cost targets from full training data into a shared source‑action scorer that re‑evaluates partial observations and source metadata after each query, applying a hard cost mask. Experiments show that O‑MPAC achieves high accuracy (0.965) in a routing task and outperforms several baselines on six real tasks, achieving the highest mean budget‑integrated accuracy on five of them.
arXiv:2606. 05606v1 Announce Type: new Abstract: LLM post-training often relies on reinforcement learning methods that sample multiple rollouts per prompt, yet most existing approaches use a fixed rollout budget for every prompt, despite large differences in the training signal different prompts provide.
arXiv:2609.38860v1 Announce Type: cross Abstract: Learning from human preferences is central to large language model (LLM) alignment, but human preference annotation is costly. Active preference lear...
arXiv:2608. 11560v1 Announce Type: new Abstract: Personalizing marketing messages with contextual multi-armed bandits (CMABs) drives real business value, yet the objective that ultimately matters - a downstream conversion - is observed only weeks later, too late to drive online learning.
arXiv:2608.24858v1 Announce Type: new Abstract: Marginalized importance weighting evaluates a target policy by reweighting offline state-action samples with its discounted occupancy ratio, characteri...
arXiv:2608. 14761v1 Announce Type: cross Abstract: At a finite public-chance cut, counterfactual regret minimization (CFR) must choose how many outcomes to evaluate before each regret update.
The paper introduces Budget-Constrained Causal Bandits (BCCB), an online framework that learns individual treatment effects, explores uncertain users, and manages budget pacing simultaneously. It derives a per-arrival decision rule from a KKT condition of a Lagrangian relaxation, providing a principled algorithmic foundation. Experiments on the Criteo Uplift dataset show BCCB outperforms offline pipelines and other online baselines, especially when historical data is scarce (below 7,500 observations).
Personalizing marketing messages with contextual multi-armed bandits (CMABs) drives real business value, yet the objective that ultimately matters - a downstream conversion - is observed only weeks later, too late to drive online learning. Teams therefore train the bandit on a fast proxy reward, and separately must judge whether a contextual bandit is worth its complexity over sending one best message.
arXiv:2606. 02595v1 Announce Type: new Abstract: Dynamic pricing in short-term rental (STR) markets presents a distinctive challenge for online learning algorithms: pricing decisions carry significant financial risk, operators require explainability, and market feedback is sparse (one booking outcome per listed night).