arXiv Machine Learning

Attribution Markets: A Fisher-Market Formulation for Fractional Credit Assignment Between Planned Tasks and Performed Actions

arXiv:2607. 20694v1 Announce Type: new Abstract: Personal and organizational planning systems maintain two records that drift apart: what was planned (a task's effort budget) and what was done (a logged action's duration and description).

arXiv Machine Learning
Jun 3

Human-in-the-Loop Contextual Bandits for Short-Term Rental Dynamic Pricing: Structural Equivalence of Historical Warm-Up and Approval-Gated Live Learning

arXiv:2606. 02595v1 Announce Type: new Abstract: Dynamic pricing in short-term rental (STR) markets presents a distinctive challenge for online learning algorithms: pricing decisions carry significant financial risk, operators require explainability, and market feedback is sparse (one booking outcome per listed night).

By Oleg Miroshnichenko
arXiv AI
Sep 15

ReCAST: Reward Credit Assignment across Timesteps for Online Diffusion Reinforcement

ReCAST is a method for assigning credit to multiple rewards during diffusion model training by using a reward-by-timestep weight matrix that respects user-specified reward budgets while ensuring equal total weight per denoising step. It allocates weight based on each reward’s informativeness, measured by its Rényi discriminability gain at each step, allowing rewards to contribute more when they are most informative. Experiments on SD3.5‑Medium with two four‑reward settings show that ReCAST improves or matches training rewards, enhances held‑out judges, and is preferred by an independent LLM‑as‑a‑Judge, indicating generalizable benefits.

By Yihang Chen, Yuanhao Ban, Kuei-Chun Kao, Cho-Jui Hsieh
arXiv AI
Aug 11

Beyond Solvability: Task Learnability as a Static Prior for LLM RL Post-Training

arXiv:2608. 09217v1 Announce Type: cross Abstract: Reinforcement learning (RL) has become a central post-training paradigm for eliciting reasoning capabilities in large language models, yet uniform task sampling allocates compute without regard to differences in how tasks respond to optimization.

By Ting Zhou, Zhenqing Ling, Daoyuan Chen, Qianli Shen, Yilun Huang, Ying Shen, Yaliang Li
arXiv Machine Learning
Jul 9

Best-Arm Identification with Generative Proxy

arXiv:2607. 06879v1 Announce Type: new Abstract: Best-arm identification is a canonical model for data-driven decision-making, but in many applications each reward observation is costly.

By Tianyi Ma, Hanzhang Qin, Ruihao Zhu, Jierui Zuo