arXiv AI By Zeyu Liu, Xuanzhi Feng, Sing Kwong Lai, Yuanchen Gao, Xiaoyi Pang, Hualei Zhang, Jingcai Guo, Jie Zhang, Song Guo

Mitigating Bias in Low-SNR Financial Reinforcement Learning via Quantum Representations

Read the original on arXiv AI →

arXiv:2606. 10448v1 Announce Type: cross Abstract: The financial market is a typical low signal-to-noise ratio (SNR) setting, which often destabilizes off-policy maximum-entropy methods like Soft Actor-Critic (SAC).

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Sep 17

APGEM: Adaptive Policy-Guided Error Mitigation for Quantum Reinforcement Learning on a Real-World CVRP Case Study

The paper introduces APGEM, an adaptive controller that dynamically selects among four error‑mitigation techniques—Zero‑Noise Extrapolation, Probabilistic Error Cancellation, Clifford Data Regression, and Readout Error Mitigation—based on a utility function and Q‑learning scores. Applied to a realistic Delhi‑based Capacitated Vehicle Routing Problem, the adaptive approach improves the quantum reinforcement learning agent’s approximation ratios from 0.84‑0.87 to 0.92‑0.94 under high noise, outperforming constructive heuristics and approaching metaheuristics. The controller’s strategy shifts from a Clifford‑data‑regression‑heavy regime early in training to a balanced use of all techniques as training progresses, demonstrating regime‑dependent selection.

By Shabir Ahmad Sofi, Bisma Majid, Mir Mohammad Yousuf
arXiv AI
Aug 3

DreamQAS: Learning a Decision-Useful World Model for VQE-Efficient Quantum Architecture Search

arXiv:2607. 29491v1 Announce Type: cross Abstract: Reinforcement-learning-based quantum architecture search (RL-QAS) repeatedly optimizes a variational quantum eigensolver (VQE) after extending a circuit, although circuit construction and action legality are deterministic and known.

By Jiayang Niu, Yan Wang, Jie Li, Ke Deng, Azadeh Alavi, Muhammad Usman, Yongli Ren
arXiv Machine Learning
Sep 1

Titans-QFWP: A Regime-Aware Hybrid Quantum Fast Weight Programmer for Portfolio Optimization

Titans-QFWP is a hybrid reinforcement learning architecture that combines a Quantum Fast Weight Programmer with a Titans-style memory system (Persistence, Surprise, and Forgetting) for adaptive portfolio optimization. It employs an enhanced A3C² framework with Hungarian-aligned K‑means clustering and scaled log‑return rewards to handle high‑dimensional market features. Tested on 468 S&P 500 stocks with about 3,000 trainable parameters, the model achieves strong performance metrics (median ARR 0.4260, Calmar 8.5504, IR 0.8427) and demonstrates that quantum gating reshapes memory roles to support drawdown control, return generation, and stabilization.

By Ming-Kai Hung, Jun-Hao Chen, Yun-Cheng Tsai, Samuel Yen-Chi Chen