arXiv Machine Learning

GRPO-QM: Target Preserving Exploration for Quantum Tomography

arXiv AI
Aug 3

DreamQAS: Learning a Decision-Useful World Model for VQE-Efficient Quantum Architecture Search

arXiv:2607. 29491v1 Announce Type: cross Abstract: Reinforcement-learning-based quantum architecture search (RL-QAS) repeatedly optimizes a variational quantum eigensolver (VQE) after extending a circuit, although circuit construction and action legality are deterministic and known.

By Jiayang Niu, Yan Wang, Jie Li, Ke Deng, Azadeh Alavi, Muhammad Usman, Yongli Ren
arXiv Machine Learning
Sep 11

Generative Replay Mitigates Sample Starvation in Quantum Architecture Search

The paper introduces GenQAS, a tensor network‑guided reinforcement learning framework that uses a learned local transition model to generate synthetic circuit transitions for prioritized generative replay. By mixing these synthetic transitions with real experience during Double Deep Q‑Network updates, GenQAS addresses sample starvation in quantum architecture search. Across benchmarks ranging from 6 to 15 qubits, the method improves success probabilities, identifies compact circuits, and reduces steps to chemical accuracy by up to 92.7%.

By Akash Kundu, Amit Kumar Jaiswal, Sebastian Feld, Prayag Tiwari
arXiv Machine Learning
Sep 21

$\lambda$-Controlled GRPO: Turning Flow-Matching Ratio Instability into a Budgeted Resource

arXiv:2609. 22041v1 Announce Type: new Abstract: Reinforcement learning is increasingly used to align image generators with reward signals, and Flow-GRPO recently extended this paradigm to flow-matching models by treating the denoising sampler as a stochastic policy that can be optimized from reward feedback.

By Yufeng Wang, Parivesh Priye, Meeshawn Marathe, Ramit Pahwa
arXiv AI
Aug 6

Agentic Reinforcement Learning with Observation-Calibrated Self-Distillation

arXiv:2608. 04788v1 Announce Type: cross Abstract: Large language model agents are commonly trained through reinforcement learning with sparse trajectory-level rewards, which offer limited guidance on how strongly individual tokens should be updated.

By Yi Yang, Cong Qin, Xiaodan Liu, Chishui Chen, Qing Dong, Yan Zhang, Cao Liu, Zhao Yang, Lu Pan, Jiaye Lin, Yi Feng
Hugging Face Trending Papers
Aug 5

Agentic Reinforcement Learning with Observation-Calibrated Self-Distillation

Large language model agents are commonly trained through reinforcement learning with sparse trajectory-level rewards, which offer limited guidance on how strongly individual tokens should be updated. On-Policy Self-Distillation (OPSD) addresses this by re-scoring generated tokens under a privileged replay view to obtain dense, token-level supervision.