Regularized Offline Policy Optimization with Posterior Hybrid Bayesian Belief
arXiv:2606. 00680v1 Announce Type: new Abstract: Offline reinforcement learning (RL) aims to optimize policies from pre-collected datasets.
arXiv:2607. 10207v1 Announce Type: cross Abstract: Data-driven optimization often requires collecting data to estimate uncertain model parameters before solving the underlying decision problem.
arXiv:2606. 00680v1 Announce Type: new Abstract: Offline reinforcement learning (RL) aims to optimize policies from pre-collected datasets.
arXiv:2407. 04900v2 Announce Type: replace Abstract: Numerous existing studies have examined the performance of Sample Average Approximation (SAA) in the fundamental newsvendor problem.
arXiv:2606. 17805v1 Announce Type: new Abstract: Data acquisition is a major bottleneck for learning in real-time streams: analysts must decide on the fly which labels to purchase while respecting a rolling budget.
The paper presents a data‑driven framework for multi‑period lost‑sales inventory control when demand is censored, meaning stockouts only reveal that demand exceeded the stocking level. It introduces a new cost decomposition for base‑stock policies and a biased sample‑average approximation (SAA) method, leading to two algorithms: an upper‑biased SAA that achieves near‑optimal sample complexity under an offline coverage condition, and a lower‑biased SAA that actively generates coverage to achieve near‑optimal online regret. The biased SAA approach offers a general principle for applying pessimism and optimism in settings with censored feedback.
arXiv:2607. 04708v1 Announce Type: cross Abstract: Agentic AI is shifting online shopping from search toward delegated purchasing, where autonomous buying agents monitor markets and decide when to buy on a consumer's behalf.
arXiv:2607. 26509v1 Announce Type: new Abstract: Deep off-policy reinforcement learning algorithms for continuous control typically rely on neural value function approximation to guide policy improvement.
arXiv:2405.08253v4 Announce Type: replace-cross Abstract: This paper develops a framework for learning in discounted infinite-horizon Markov decision processes (MDPs) with Borel state and action spac...
arXiv:2511.20413v2 Announce Type: replace-cross Abstract: \emph{Decision-focused learning} (DFL) trains predictive models to optimize downstream decisions rather than prediction accuracy alone. While...
arXiv:2608. 01151v1 Announce Type: cross Abstract: In this paper, we consider stochastic optimal control problems with infinite-horizon joint chance constraints.
The paper presents a large deviations framework for efficient data acquisition in infinite-horizon reinforcement learning, introducing the exponential decay rate of policy-selection error probability as a key efficiency metric. It derives a variational characterization leading to a nested optimization problem, then proposes a tractable convex relaxation and a lazy one-step projected subgradient method to construct an adaptive data acquisition policy. The resulting algorithm is shown to be near-robustly optimal under the proposed criterion, with extensions to linear function approximation and supporting numerical experiments.
arXiv:2607. 10694v1 Announce Type: cross Abstract: We study the problem of optimal continual fine-tuning for a pre-trained Foundation Model deployed at a resource-limited device.
arXiv:2606. 05606v1 Announce Type: new Abstract: LLM post-training often relies on reinforcement learning methods that sample multiple rollouts per prompt, yet most existing approaches use a fixed rollout budget for every prompt, despite large differences in the training signal different prompts provide.