arXiv:2606. 06744v1 Announce Type: new Abstract: Two-sided matching markets often involve information that unfolds over time through interviews, repeated interaction, learning, and separation.
By Haijing Zong, Yancheng Liang, Boyang Zhou, Natasha Jaques
arXiv:2608. 09130v1 Announce Type: cross Abstract: Allocating limited computation among concurrent learning tasks is difficult when each task must reach a target loss before a deadline but its required training effort is unknown.
By Hanye Zhao, Muning Wen, Yong Yu, Weinan Zhang
arXiv:2606. 00367v1 Announce Type: cross Abstract: Reinforcement learning problems typically define the goal as maximizing the expected value of a scalar reward function.
By Jonathan Cola\c{c}o Carr, Prakash Panangaden, Doina Precup, Benjamin Van Roy
The paper introduces Graph-Guided Quasimetric Dense Reward (G2QDR), a framework that learns a state connectivity model to predict pairwise connectivity strengths in asymmetric environments. These strengths are converted into scalar auxiliary dense rewards, offering continuous guidance across hierarchical levels. G2QDR can be integrated into any existing Goal-Conditioned Hierarchical Reinforcement Learning architecture and shows empirical performance improvements in sparse reward settings with modest computational cost.
By Shuyuan Zhang, Zihan Wang, Xiao-Wen Chang, Doina Precup
arXiv:2608.23664v1 Announce Type: cross
Abstract: Reward fine-tuning is becoming an important tool for adapting diffusion models to human preferences and task-specific objectives, but existing method...
By Jaemoo Choi, Wei Guo, Yuchen Zhu, Arash Vahdat, Molei Tao, Julius Berner, Yongxin Chen
TIGPO (Temporal Instance-Graph Policy Optimization) extends graph-based credit assignment for long-horizon LLM agents by maintaining a persistent transition graph per task across policy updates. It allocates rollout budgets to both new exploration and revisiting past tasks, pairing current rollouts with earlier ones to create cross‑temporal references that stabilize advantage estimation. Experiments on ALFWorld and WebShop show TIGPO consistently outperforms previous group‑based and graph‑based policy optimization methods.
By Jinwei Gan