arXiv:2606. 00680v1 Announce Type: new Abstract: Offline reinforcement learning (RL) aims to optimize policies from pre-collected datasets.
By Hongqiang Lin, Pengfei Wang, Nenggan Zheng
arXiv:2407. 04900v2 Announce Type: replace Abstract: Numerous existing studies have examined the performance of Sample Average Approximation (SAA) in the fundamental newsvendor problem.
By Jiameng Lyu, Shilin Yuan, Bingkun Zhou, Yuan Zhou
arXiv:2606. 17805v1 Announce Type: new Abstract: Data acquisition is a major bottleneck for learning in real-time streams: analysts must decide on the fly which labels to purchase while respecting a rolling budget.
By Xiwen Huang, Pierre Pinson
The paper presents a data‑driven framework for multi‑period lost‑sales inventory control when demand is censored, meaning stockouts only reveal that demand exceeded the stocking level. It introduces a new cost decomposition for base‑stock policies and a biased sample‑average approximation (SAA) method, leading to two algorithms: an upper‑biased SAA that achieves near‑optimal sample complexity under an offline coverage condition, and a lower‑biased SAA that actively generates coverage to achieve near‑optimal online regret. The biased SAA approach offers a general principle for applying pessimism and optimism in settings with censored feedback.
By Yuxuan Han, Xiaoyu Fan, Jiawei Zhang, Zhengyuan Zhou
arXiv:2607. 04708v1 Announce Type: cross Abstract: Agentic AI is shifting online shopping from search toward delegated purchasing, where autonomous buying agents monitor markets and decide when to buy on a consumer's behalf.
By Mingyang Fu, Ming Hu
arXiv:2607. 26509v1 Announce Type: new Abstract: Deep off-policy reinforcement learning algorithms for continuous control typically rely on neural value function approximation to guide policy improvement.
By Gong Gao, Xiao Lai, Ziqi Xie, Guojie Chen, Xianhui Liu, Weidong Zhao