The paper introduces Quantum‑Harbor, a virtual laboratory that lets AI agents interact with quantum systems in a controlled setting, enabling verification of their actions and conclusions. Using this platform, the authors created QIQCBench, a benchmark of 49 expert‑authored tasks covering calibration, control, error correction, compilation, sensing, and networking. Testing 17 state‑of‑the‑art agentic systems on QIQCBench revealed wide variation in verified performance, highlighting a gap between demonstrated capability and reliable operation and positioning Quantum‑Harbor as a foundation for measuring progress toward verified autonomy in quantum engineering.
By Naixu Guo, Changhao Li, Siyu Cheng, Qicheng Tang, Binzhao Luo, Bikun Li, Yuxuan Du, Shihao Ru, Jiaqi Cai
arXiv:2508. 20134v2 Announce Type: replace Abstract: Programming quantum circuits at the OpenQASM level is essential for achieving hardware-aware optimization and reliable execution on noisy intermediate-scale quantum (NISQ) devices, yet it remains challenging due to the need for domain-specific planning, iterative code synthesis, and low-level calibration.
By Zhenxiao Fu, Lei Jiang, Yilun Xu, Gang Huang, Fan Chen
arXiv:2607. 25865v1 Announce Type: cross Abstract: Quantum error correction (QEC) is indispensable for scalable fault-tolerant quantum computing.
By Ge Yan, Shanchuan Li, Pengyue Ma, Qixin Zhang, Pingchuan Ma, Jianping Wang, Min-Hsiu Hsieh, Yuxuan Du
arXiv:2607. 25834v1 Announce Type: cross Abstract: Quantum computers are moving from research laboratories to industrial machines accessible via the cloud and integrated into high-performance computing facilities.
By Constantin Dalyac, Alexandre Dauphin, Lo\"ic Henriet, Christophe Jurczak
QEncodeBench evaluates whether large language models can translate classical constraint problems into verified quantum phase oracles. The benchmark measures the correctness of generated circuits using an adversarial self‑validated verifier that checks full solution‑set equivalence while enforcing resource limits. Results show that models lacking a reasoning mode perform poorly, whereas enabling native reasoning improves accuracy tenfold; semantic errors dominate, and neuro‑symbolic pipelines close most gaps by delegating critical composition to deterministic procedures.
By Xujun Che, Hanhan Wu, Yuchen Yuan, Chenyang Yu
arXiv:2607. 25145v1 Announce Type: cross Abstract: We implement an agentic AI workflow built around a large language model (LLM) agent for autonomous experiments with nitrogen-vacancy (NV) centers in diamond.
By Takuya Isogawa, Ryotaro Okabe, Nutdech Phadetsuwannukun, Mingda Li, Paola Cappellaro
arXiv:2605. 30358v2 Announce Type: replace Abstract: Quantum computing remains in the Noisy Intermediate-Scale Quantum (NISQ) era, with performance constrained by noise.
By Zhenxiao Fu, Lei Jiang, Fan Chen
QART is a quantum‑classical hybrid architecture that augments a language model with quantum encoding, CIM‑based QUBO optimization, and quantum decoding to improve long‑horizon reasoning. The authors claim that, under certain assumptions, QART can maintain a non‑zero probability of recovering an optimal reasoning path while traditional autoregressive LLMs see their acceptance probability drop to zero as cumulative risk grows. Experiments on six benchmarks with three backbone models show that QART outperforms the baselines in 14 of 15 pairings, with relative gains up to 84.0% on SciCode.
By Lehao Lin, Yuheng Cheng, Guolong Liu, Yao Li, Xuning Tan, Xiyuan Zhou, Ruixi Zou, Shi Wang, Huan Zhao, Wenxuan Liu, Haifeng Wu, Junhua Zhao
arXiv:2603. 13191v2 Announce Type: replace-cross Abstract: While large language models (LLMs) have transformed AI agents into proficient executors of computational materials science, performing a hundred simulations does not make a researcher.
By Haonan Huang
arXiv:2603.17043v2 Announce Type: replace
Abstract: Physics-aware multimodal large language models (MLLMs) can localize and characterize two-dimensional (2D) material flakes from optical microscopy i...
By Sankalp Pandey, Thanh-Dat Truong, Xuan-Bac Nguyen, Hoang-Quan Nguyen, Tim Faltermeier, Nicholas Borys, Hugh Churchill, Khoa Luu
arXiv:2507. 18606v2 Announce Type: replace-cross Abstract: Reinforcement learning (RL) provides a principled framework for decision-making in partially observable environments, which can be modeled as Markov decision processes and compactly represented through dynamic decision Bayesian networks.
By Gilberto Cunha, Alexandra Ram\^oa, Andr\'e Sequeira, Michael de Oliveira, Lu\'is Barbosa
arXiv:2609.05842v1 Announce Type: cross
Abstract: Reinforcement learning with verifiable rewards enables large language models to think slowly, but the same training can induce policy collapse: proba...
By Xiansheng Cai, Xiu-Hao Deng, Kun Chen