← Back to all news
arXiv Machine Learning October 8, 2026 By Zhaohua Chen

Online Resource Allocation with an Endogenous Markov State: Fewer LP Solves Earn More

Read the original on arXiv Machine Learning →

The Flow has not summarised this story yet — read it at arXiv Machine Learning.

  • benchmarks

One email a morning, machine-written

One email a day, machine-written, one click to leave. We never share your address.

Related stories

arXiv Machine Learning
Jul 3

Online Resource Allocation with Continuous Random Consumption: Regret under Degeneracy

arXiv:2607. 02196v1 Announce Type: new Abstract: We study online resource allocation when both rewards and consumption sizes may be continuously distributed.

By Jiawei Zhang
More like this →
arXiv Machine Learning
Sep 30

Nonpreemptive Scheduling While Learning Context-Dependent Service Rates

arXiv:2609.37660v1 Announce Type: new Abstract: We study nonpreemptive contextual queueing bandits in a single-server system. Each job is represented by a $d$-dimensional context vector; in each roun...

By Wansoo Choi, Seoungbin Bae, Dabeen Lee
More like this →
arXiv Statistics ML
Sep 2

Model Predictive Control is almost Optimal for Heterogeneous Restless Multi-armed Bandits

arXiv:2511.08097v2 Announce Type: replace-cross Abstract: We consider a general infinite horizon Heterogeneous Restless multi-armed Bandit (RMAB). Heterogeneity is a fundamental problem for many real...

By Dheeraj Narasimha, Nicolas Gast
reinforcement-learning
More like this →
arXiv Machine Learning
Jul 23

Asymptotically Optimal Regret for Reinforcement Learning without Horizon Dependence

arXiv:2607. 19854v1 Announce Type: new Abstract: We study horizon-free regret minimization for finite-horizon time-homogeneous tabular Markov decision processes with $S$ states, $A$ actions, horizon $H$, and per-trajectory total reward bounded by $1$.

By Runlong Zhou, Zihan Zhang, Maryam Fazel, Simon S. Du
reinforcement-learning
More like this →
arXiv Machine Learning
Sep 30

Fine-Tuning on Self-Generated and Reward-Weighted Data: Learning Dynamics, Convergence Rates, and Benefits of Off-Policyness

arXiv:2609.36945v1 Announce Type: new Abstract: We study the learning dynamics of fine-tuning a policy model on self-generated and reward-weighted data, with particular focus on a generalized version...

By Zhiwei Wang, Yanxi Chen, Yaliang Li, Bolin Ding
llmsreinforcement-learningfine-tuning
More like this →
arXiv Machine Learning
Jun 29

Learning in Markovian bandits with non-observable states and constrained decision epochs

arXiv:2606. 27448v1 Announce Type: new Abstract: This paper studies the problem of regret minimization in Markovian bandits with \emph{non-observable states} and possibly \emph{constrained} decision epochs.

By Thomas Hira, Victor Boone, Urtzi Ayesta, Ina Maria Verloop
reinforcement-learningbenchmarkssafety
More like this →
About Pricing API Newsletter Sources Privacy Terms Refunds Accessibility Provider info Contact RSS

The Flow links to publishers and never republishes their articles. Summaries are machine-generated.

v1.1.0 · 5f852ea