arXiv:2609.14711v2 Announce Type: replace
Abstract: Bayesian quantum tomography requires efficient inference while preserving a posterior fixed by the prior and Born likelihood. Learned transport pro...
By Yufeng Wang, Parivesh Priye, Lu Wei, Haibin Ling
arXiv:2607. 29491v1 Announce Type: cross Abstract: Reinforcement-learning-based quantum architecture search (RL-QAS) repeatedly optimizes a variational quantum eigensolver (VQE) after extending a circuit, although circuit construction and action legality are deterministic and known.
By Jiayang Niu, Yan Wang, Jie Li, Ke Deng, Azadeh Alavi, Muhammad Usman, Yongli Ren
The paper introduces GenQAS, a tensor network‑guided reinforcement learning framework that uses a learned local transition model to generate synthetic circuit transitions for prioritized generative replay. By mixing these synthetic transitions with real experience during Double Deep Q‑Network updates, GenQAS addresses sample starvation in quantum architecture search. Across benchmarks ranging from 6 to 15 qubits, the method improves success probabilities, identifies compact circuits, and reduces steps to chemical accuracy by up to 92.7%.
By Akash Kundu, Amit Kumar Jaiswal, Sebastian Feld, Prayag Tiwari
arXiv:2609. 22041v1 Announce Type: new Abstract: Reinforcement learning is increasingly used to align image generators with reward signals, and Flow-GRPO recently extended this paradigm to flow-matching models by treating the denoising sampler as a stochastic policy that can be optimized from reward feedback.
By Yufeng Wang, Parivesh Priye, Meeshawn Marathe, Ramit Pahwa
arXiv:2608. 04788v1 Announce Type: cross Abstract: Large language model agents are commonly trained through reinforcement learning with sparse trajectory-level rewards, which offer limited guidance on how strongly individual tokens should be updated.
By Yi Yang, Cong Qin, Xiaodan Liu, Chishui Chen, Qing Dong, Yan Zhang, Cao Liu, Zhao Yang, Lu Pan, Jiaye Lin, Yi Feng
arXiv:2609.38329v1 Announce Type: new
Abstract: Group-relative RL methods such as Flow-GRPO post-train image generators by exploring with isotropic Gaussian noise added at every denoising step. This...
By Shuyue Stella Li, Xiaochuang Han, Yulia Tsvetkov, Luke Zettlemoyer
arXiv:2609.05842v1 Announce Type: cross
Abstract: Reinforcement learning with verifiable rewards enables large language models to think slowly, but the same training can induce policy collapse: proba...
By Xiansheng Cai, Xiu-Hao Deng, Kun Chen
arXiv:2607. 28916v1 Announce Type: cross Abstract: Multistep credit assignment is critical for sample-efficient reinforcement learning, yet managing off-policy bias in Q-learning remains a fundamental challenge.
By Brett Daley
arXiv:2608.21595v1 Announce Type: new
Abstract: Reinforcement learning with verifiable rewards (RLVR) improves the reasoning ability of vision-language models (VLMs), and diversifying the rollouts wi...
By Michael Jerge, Joseph Pelczar, Justin Downes
Large language model agents are commonly trained through reinforcement learning with sparse trajectory-level rewards, which offer limited guidance on how strongly individual tokens should be updated. On-Policy Self-Distillation (OPSD) addresses this by re-scoring generated tokens under a privileged replay view to obtain dense, token-level supervision.
arXiv:2608.21946v1 Announce Type: cross
Abstract: Reinforcement learning with outcome-based objectives such as GRPO enables LLM-based agents to solve complex, long-horizon tasks, yet the reusable exp...
By Can Xie, Yuyi Zhou, Wen Yang, Ziyi zhang, Siyao Song, Yingzhuo Deng, Shuo Ren, Jiajun Zhang
arXiv:2608. 01597v1 Announce Type: new Abstract: Search-augmented LM agents are typically trained with a binary exact-match reward, which throws away most of what a failed trajectory tells us about why it failed.
By Haowei Liu, Jiamian Wang, Hsin-Tai Wu, Zhiqiang Tao, Yi Fang