arXiv AI

Benchmarking API Drift in LLM-Generated Quantum Code Across Successive SDK Versions

arXiv:2607. 04072v1 Announce Type: cross Abstract: Large language models can generate plausible quantum code, but it is unclear whether they can reliably target the specific software development kit (SDK) version requested by the user.

arXiv Machine Learning
Sep 7

Impact of Data Loss in Postprocessing on Training and Inference of Quantum Neural Networks

The paper investigates how postprocessing routines in quantum neural network software can cause significant data loss when run on large quantum hardware. In a case study of Qiskit’s “SamplerQNN”, a filter that assumes measurement bit‑strings are in virtual qubit space removed 85–99.6% of valid shots on IBM backends, leading to unnormalised probability vectors and distorted predictions. This loss caused inference accuracy to drop from 0.94 to 0.39 and compressed training loss signals by 22–27×, severely reducing optimizer sensitivity. The authors implemented a layout‑based marginalisation fix that was merged into the library to make “SamplerQNN” forward‑compatible with current and future hardware.

By Soraya V. Panambalom, Edoardo Altamura, Nick Chancellor, Jonte R. Hance
Hugging Face Trending Papers
Sep 4

Impact of Data Loss in Postprocessing on Training and Inference of Quantum Neural Networks

The paper investigates how postprocessing routines in quantum neural network software can inadvertently discard a large portion of valid measurement data when run on real quantum hardware. In a case study of Qiskit’s “SamplerQNN”, a filter that assumes virtual qubit space caused 85–99.6% of measurement shots to be lost on IBM backends, leading to unnormalised probability vectors, degraded inference accuracy (from 0.94 to 0.39), and a 22–27× compression of the training loss signal. The authors provide a layout‑based marginalisation fix that has been merged into the library to ensure forward‑compatibility with current and future hardware.

arXiv AI
Jul 24

PennySynth: RAG-Driven Data Synthesis for Automated Quantum Code Generation

arXiv:2605. 25572v2 Announce Type: replace-cross Abstract: The growing complexity of quantum programming frameworks has exposed a critical limitation in existing large language model (LLM)-based code assistants: general-purpose models hallucinate PennyLane-specific gate names, misplace device configurations, and produce structurally invalid circuits when faced with specialized quantum coding challenges.

By Minghao Shao, Nouhaila Innan, Hariharan Janardhanan, Muhammad Kashif, Alberto Marchisio, Muhammad Shafique
Hugging Face Trending Papers
5d ago

QC-Stark: A Multi-Task Benchmark Revealing Capability Dissociations in LLMs Evaluated on Quantum Computing Tasks

QC-Stark is a benchmark that evaluates large language models on 11 quantum computing tasks, including circuit construction, debugging, compilation, error correction, and simulation. It comprises 2,750 evaluations across 10 models, 5 difficulty levels, and 5 seeds, revealing that overall model rankings can hide significant per-task differences, with Spearman correlation being insignificant for 4 of the 11 tasks. The benchmark’s measurement quality is validated by a 2‑parameter Item Response Theory model, prompt sensitivity analysis confirms ranking robustness, and all tasks are auto‑verifiable via execution, eliminating the need for manual evaluation. The code and data are publicly available on Hugging Face.

arXiv AI
Jul 13

QAgent: An LLM-based Multi-Agent System for Autonomous OpenQASM programming

arXiv:2508. 20134v2 Announce Type: replace Abstract: Programming quantum circuits at the OpenQASM level is essential for achieving hardware-aware optimization and reliable execution on noisy intermediate-scale quantum (NISQ) devices, yet it remains challenging due to the need for domain-specific planning, iterative code synthesis, and low-level calibration.

By Zhenxiao Fu, Lei Jiang, Yilun Xu, Gang Huang, Fan Chen
arXiv AI
Sep 7

Qlippy: A Retrieval-Augmented GenAI Assistant for Reproducible Quantum Workflows and Experiment Tracking

Qlippy is a retrieval‑augmented generative AI assistant designed to support reproducible quantum software development. It is embedded in the development environment and grounds its responses in a curated corpus of quantum‑software‑engineering knowledge, explaining reproducibility and provenance concepts in context. The assistant augments Qiskit programs with MLflow‑based experiment tracking that follows the QProv schema, thereby reducing reliance on large language models and enabling low‑cost, privacy‑preserving local deployment.

By Mahee Gamage, Vlad Stirbu
arXiv Machine Learning
Jul 15

VQCSim: When Does Compile-Once Statevector Simulation Beat Generic Quantum Frameworks?

arXiv:2607. 11985v1 Announce Type: cross Abstract: Hybrid quantum-classical machine learning workflows repeatedly evaluate many small parametrized circuits during training and model exploration.

By Anton Firc, Martin Pere\v{s}\'ini, Vojt\v{e}ch Mr\'azek, Kamil Malinka, Vojt\v{e}ch Stan\v{e}k, Zbyn\v{e}k Li\v{c}ka, Nouhaila Innan, Walid El Maouaki, Alberto Marchisio, Muhammad Shafique
arXiv AI
Sep 16

QART: A Quantum-Classical Hybrid Architecture for Long-Horizon Reasoning -- Exploring a Conditional Path toward Quantum Scaling

QART is a quantum‑classical hybrid architecture that augments a language model with quantum encoding, CIM‑based QUBO optimization, and quantum decoding to improve long‑horizon reasoning. The authors claim that, under certain assumptions, QART can maintain a non‑zero probability of recovering an optimal reasoning path while traditional autoregressive LLMs see their acceptance probability drop to zero as cumulative risk grows. Experiments on six benchmarks with three backbone models show that QART outperforms the baselines in 14 of 15 pairings, with relative gains up to 84.0% on SciCode.

By Lehao Lin, Yuheng Cheng, Guolong Liu, Yao Li, Xuning Tan, Xiyuan Zhou, Ruixi Zou, Shi Wang, Huan Zhao, Wenxuan Liu, Haifeng Wu, Junhua Zhao
arXiv Machine Learning
Sep 18

QEncodeBench: Can Large Language Models Encode Classical Problems into Verified Quantum Oracles?

QEncodeBench evaluates whether large language models can translate classical constraint problems into verified quantum phase oracles. The benchmark measures the correctness of generated circuits using an adversarial self‑validated verifier that checks full solution‑set equivalence while enforcing resource limits. Results show that models lacking a reasoning mode perform poorly, whereas enabling native reasoning improves accuracy tenfold; semantic errors dominate, and neuro‑symbolic pipelines close most gaps by delegating critical composition to deterministic procedures.

By Xujun Che, Hanhan Wu, Yuchen Yuan, Chenyang Yu