arXiv AI

Agentic AI for Scientific Reasoning in Autonomous Quantum Sensing Experiments

arXiv:2607. 25145v1 Announce Type: cross Abstract: We implement an agentic AI workflow built around a large language model (LLM) agent for autonomous experiments with nitrogen-vacancy (NV) centers in diamond.

arXiv AI
Aug 11

QuantumMind: Constraint-Grounded Agentic Reasoning for Speedup Analysis in Quantum Computing

arXiv:2608. 07743v1 Announce Type: new Abstract: Identifying a meaningful quantum speedup requires more than matching a classical problem to a familiar quantum primitive: the claim must preserve the task, respect access and output models, expose required promises, and remain within a defensible complexity scope.

By Yijing Zuo, Zhe Fu, Zihan Nie, Zhihui Zhu, Haohan Wang
arXiv AI
Jul 13

QAgent: An LLM-based Multi-Agent System for Autonomous OpenQASM programming

arXiv:2508. 20134v2 Announce Type: replace Abstract: Programming quantum circuits at the OpenQASM level is essential for achieving hardware-aware optimization and reliable execution on noisy intermediate-scale quantum (NISQ) devices, yet it remains challenging due to the need for domain-specific planning, iterative code synthesis, and low-level calibration.

By Zhenxiao Fu, Lei Jiang, Yilun Xu, Gang Huang, Fan Chen
arXiv AI
Sep 16

Evaluating Verified Autonomy in Quantum Engineering

The paper introduces Quantum‑Harbor, a virtual laboratory that lets AI agents interact with quantum systems in a controlled setting, enabling verification of their actions and conclusions. Using this platform, the authors created QIQCBench, a benchmark of 49 expert‑authored tasks covering calibration, control, error correction, compilation, sensing, and networking. Testing 17 state‑of‑the‑art agentic systems on QIQCBench revealed wide variation in verified performance, highlighting a gap between demonstrated capability and reliable operation and positioning Quantum‑Harbor as a foundation for measuring progress toward verified autonomy in quantum engineering.

By Naixu Guo, Changhao Li, Siyu Cheng, Qicheng Tang, Binzhao Luo, Bikun Li, Yuxuan Du, Shihao Ru, Jiaqi Cai
arXiv Machine Learning
Sep 18

QEncodeBench: Can Large Language Models Encode Classical Problems into Verified Quantum Oracles?

QEncodeBench evaluates whether large language models can translate classical constraint problems into verified quantum phase oracles. The benchmark measures the correctness of generated circuits using an adversarial self‑validated verifier that checks full solution‑set equivalence while enforcing resource limits. Results show that models lacking a reasoning mode perform poorly, whereas enabling native reasoning improves accuracy tenfold; semantic errors dominate, and neuro‑symbolic pipelines close most gaps by delegating critical composition to deterministic procedures.

By Xujun Che, Hanhan Wu, Yuchen Yuan, Chenyang Yu
arXiv AI
Sep 7

La Agente \'Optima: Towards Agentic Self-Driving Laboratories

La Agente ’Optima is an agentic framework that builds and manages Bayesian optimization campaigns for self‑driving laboratories, separating large language model reasoning from campaign execution. It maintains a persistent optimization state, allowing consistent repetitive loops and auditable decisions, and only returns control to the agent when interpretation or revision is needed. In tests on digital discovery tasks and physical platforms, it corrected measurement failures, improved yields, and recommended formulation changes, outperforming human‑directed campaigns in cost and material usage.

By Marcel M\"uller, Jiaru Bai, Willi Gottstein, Abhijoy Mandal, Mohammad Nazeri, Elia Savino, Yanlin Fang, Sujoy Das, Sergio Pablo Garc\'ia Carrillo, Yeonghun Kang, Juan B. P\'erez-S\'anchez, Simone Pilon, Martin Fitzner, Timothy No\"el, Frank Gu, Varinia Bernales, Al\'an Aspuru-Guzik
arXiv AI
Jun 26

Life After Benchmark Saturation: A Case Study of CORE-Bench

arXiv:2606. 26158v1 Announce Type: new Abstract: When a benchmark's accuracy saturates, it is often retired and replaced with a more challenging version.

By Nitya Nadgir, Sayash Kapoor, Kangheng Liu, Peter Kirgis, Matilda Orona, Stephan Rabanser, Tilman Bayer, Abhishek Shetty, Yue Ling, Derrick Chan-Sew, Rumi Nakagawa, Saiteja Utpala, Zachary S. Siegel, Arvind Narayanan
arXiv AI
Aug 25

The Greatness of Science Cannot Be Planned: Agentic Auto-Research is Fuzz Testing

The article argues that agentic auto‑research should be guided by dense, intermediate signals of epistemic progress rather than by sparse final benchmarks. It compares this approach to fuzz testing, where coverage provides continuous feedback that directs input mutation. The authors propose controlled experiments to test whether such signals improve discovery efficiency and reduce false positives, and demonstrate in a simulated physics setting that an AI agent using feedback‑driven search uncovers a hidden law while optimization‑driven baselines fail.

By Yifeng He, Jicheng Wang, Yinzhe Zhao, Chengyang Shi, Jiachen Liu, Hao Chen
arXiv AI
Jul 14

FIRE-Bench: Evaluating AI Agents on the Rediscovery of Scientific Insights

arXiv:2602. 02905v2 Announce Type: replace Abstract: Autonomous agents powered by large language models (LLMs) promise to accelerate scientific discovery end-to-end, but rigorously evaluating their capacity for verifiable discovery remains a central challenge.

By Zhen Wang, Fan Bai, Zhongyan Luo, Jinyan Su, Kaiser Sun, Xinle Yu, Jieyuan Liu, Kun Zhou, Claire Cardie, Mark Dredze, Zhiting Hu, Eric P. Xing