The paper introduces Quantum‑Harbor, a virtual laboratory that lets AI agents interact with quantum systems in a controlled setting, enabling verification of their actions and conclusions. Using this platform, the authors created QIQCBench, a benchmark of 49 expert‑authored tasks covering calibration, control, error correction, compilation, sensing, and networking. Testing 17 state‑of‑the‑art agentic systems on QIQCBench revealed wide variation in verified performance, highlighting a gap between demonstrated capability and reliable operation and positioning Quantum‑Harbor as a foundation for measuring progress toward verified autonomy in quantum engineering.
By Naixu Guo, Changhao Li, Siyu Cheng, Qicheng Tang, Binzhao Luo, Bikun Li, Yuxuan Du, Shihao Ru, Jiaqi Cai
arXiv:2508. 20134v2 Announce Type: replace Abstract: Programming quantum circuits at the OpenQASM level is essential for achieving hardware-aware optimization and reliable execution on noisy intermediate-scale quantum (NISQ) devices, yet it remains challenging due to the need for domain-specific planning, iterative code synthesis, and low-level calibration.
By Zhenxiao Fu, Lei Jiang, Yilun Xu, Gang Huang, Fan Chen
arXiv:2607. 25865v1 Announce Type: cross Abstract: Quantum error correction (QEC) is indispensable for scalable fault-tolerant quantum computing.
By Ge Yan, Shanchuan Li, Pengyue Ma, Qixin Zhang, Pingchuan Ma, Jianping Wang, Min-Hsiu Hsieh, Yuxuan Du
arXiv:2607. 25834v1 Announce Type: cross Abstract: Quantum computers are moving from research laboratories to industrial machines accessible via the cloud and integrated into high-performance computing facilities.
By Constantin Dalyac, Alexandre Dauphin, Lo\"ic Henriet, Christophe Jurczak
QEncodeBench evaluates whether large language models can translate classical constraint problems into verified quantum phase oracles. The benchmark measures the correctness of generated circuits using an adversarial self‑validated verifier that checks full solution‑set equivalence while enforcing resource limits. Results show that models lacking a reasoning mode perform poorly, whereas enabling native reasoning improves accuracy tenfold; semantic errors dominate, and neuro‑symbolic pipelines close most gaps by delegating critical composition to deterministic procedures.
By Xujun Che, Hanhan Wu, Yuchen Yuan, Chenyang Yu
arXiv:2607. 25145v1 Announce Type: cross Abstract: We implement an agentic AI workflow built around a large language model (LLM) agent for autonomous experiments with nitrogen-vacancy (NV) centers in diamond.
By Takuya Isogawa, Ryotaro Okabe, Nutdech Phadetsuwannukun, Mingda Li, Paola Cappellaro